Manager, AI Engineering (Tester )
Core
Lead AI quality engineering efforts for Mastercard's Operational Intelligence Program, ensuring Generative AI, LLM, and agentic systems are accurate, safe, and enterprise-ready.
Role type
Manager, AI Testing & Quality Engineering
Builds
Evaluation frameworks, test suites, red-teaming libraries, and production monitoring pipelines for Gen AI systems.
Domain
Financial Technology / Generative AI / LLM Operations
Deliverable
production ML models | infrastructure
Required skills
LLM evaluation frameworks (RAGAS, DeepEval, TruLens), Python programming, SQL, CI/CD integration, cloud AI infrastructure (AWS), observability tooling (Grafana, Datadog), agentic workflow validation, prompt injection testing, bias auditing.
Preferred skills
Experience with LangChain/LangGraph, MLOps/LLMOps pipelines (MLflow, Databricks), frontier evaluation benchmarks (MMLU, TruthfulQA).
Technologies
RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, Databricks, AWS, Grafana, Datadog, CloudWatch, LangChain, LangGraph, MLflow, SageMaker, OpenAI, Anthropic, Hugging Face.
Responsibilities
Design end-to-end LLM evaluation frameworks including automated prompt regression and hallucination detection; Build comprehensive test suites for agentic AI systems validating tool selection and inter-agent coordination; Lead structured red-teaming and adversarial testing exercises targeting prompt injection and data leakage; Execute fairness, bias, and Responsible AI audits; Design and run inference performance benchmarks measuring latency and throughput; Build production monitoring and drift detection pipelines for semantic output and embedding shifts; Define the AI testing roadmap and quality standards for the program.
Seniority
Manager, hands-on technical leadership