CareerPlanSign in

Manager, AI Engineering (Tester )

O'Fallon, Missouri💼 Full-time💰 $140,000–$140,000🗓 2026-07-24 → 2026-09-26

Core

Lead AI quality engineering efforts for Mastercard's Operational Intelligence Program, ensuring Generative AI, LLM, and agentic systems are accurate, safe, and enterprise-ready.

Role type

Manager, AI Testing & Quality Engineering

Builds

Evaluation frameworks, test suites, red-teaming libraries, and production monitoring pipelines for Gen AI systems.

Domain

Financial Technology / Generative AI / LLM Operations

Deliverable

production ML models | infrastructure

Required skills

LLM evaluation frameworks (RAGAS, DeepEval, TruLens), Python programming, SQL, CI/CD integration, cloud AI infrastructure (AWS), observability tooling (Grafana, Datadog), agentic workflow validation, prompt injection testing, bias auditing.

Preferred skills

Experience with LangChain/LangGraph, MLOps/LLMOps pipelines (MLflow, Databricks), frontier evaluation benchmarks (MMLU, TruthfulQA).

Technologies

RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, Databricks, AWS, Grafana, Datadog, CloudWatch, LangChain, LangGraph, MLflow, SageMaker, OpenAI, Anthropic, Hugging Face.

Responsibilities

Design end-to-end LLM evaluation frameworks including automated prompt regression and hallucination detection; Build comprehensive test suites for agentic AI systems validating tool selection and inter-agent coordination; Lead structured red-teaming and adversarial testing exercises targeting prompt injection and data leakage; Execute fairness, bias, and Responsible AI audits; Design and run inference performance benchmarks measuring latency and throughput; Build production monitoring and drift detection pipelines for semantic output and embedding shifts; Define the AI testing roadmap and quality standards for the program.

Seniority

Manager, hands-on technical leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.