AI Agent Quality Engineer
Core
Lead the charge in testing, stress-testing, and breaking AI agents before production to ensure they are secure, resilient, and compliant.
Role type
Senior IC AI Quality & Red Team Engineer
Builds
Automated evaluation suites, adversarial test scenarios, and guardrail catalogs for an AI agent fleet
Domain
Cloud Security / AI Safety / LLM Red Teaming
Deliverable
production ML models
Required skills
Python, REST APIs, CI/CD integration, adversarial testing frameworks, prompt injection design, drift tracking, audit evidence generation
Preferred skills
Data sensitivity classification, regulated environment audit experience, multi-attempt attack methodology, agent architecture familiarity
Technologies
DeepTeam, Garak, PyRIT, Promptfoo
Responsibilities
Build and grow automated evaluation suites for unattended agent fleet testing; Design and automate adversarial test scenarios including prompt injections and sycophancy checks; Own the 'break it on purpose' pass to extract unauthorized data or force out-of-bounds actions; Partner with Data Stewards to align test scenarios with data sensitivity; Define pass/fail thresholds for agent capabilities; Maintain guardrail and negative-test catalogs; Produce audit-ready evidence automatically; Track fleet-wide drift over time
Seniority
Senior, hands-on IC