Senior AI QA Test Automation Engineer – Agent
Core
Design and build scalable AI evaluation systems and QA tooling to validate the reliability, accuracy, and safety of a multi-agent customer-support system (CX Agent Suite) built on AWS Bedrock AgentCore.
Role type
Senior Individual Contributor (IC) AI QA Test Automation Engineer
Builds
Production-grade AI evaluation pipelines, automated testing frameworks, and quality analytics dashboards for the CX Agent Suite
Domain
Artificial Intelligence / Machine Learning / Customer Experience (CX)
Deliverable
production ML models | product features
Required skills
Python, AWS Bedrock, LLM evaluation techniques (LLM-as-a-judge, human-in-the-loop), DeepEval, agent-orchestration patterns, CI/CD integration, Kubernetes/EKS, vector databases, MLOps
Preferred skills
Strands Agents, autonomous QA agents, LLM observability (LangFuse, LangSmith), AI ethics and bias detection, gRPC/WebSockets, multi-cloud AI platforms (GCP Vertex AI, Azure AI)
Technologies
Python, AWS Bedrock, DeepEval, Strands Agents, GitLab CI/CD, Terraform, EKS, Docker, Kubernetes, OpenTelemetry, CloudWatch, X-Ray, LangFuse, Vector Databases
Responsibilities
Architect scalable AI evaluation systems and establish technical standards for QA; Design and build AI evaluation pipelines to assess multi-agent responses for accuracy, relevance, tone, hallucination, and safety; Evaluate and prototype emerging agent-orchestration frameworks (e.g., Strands Agents) for autonomous QA; Integrate AI-driven testing into enterprise CI/CD pipelines; Develop data-driven quality analytics and root-cause analysis; Mentor engineers on AI testing methodologies and architectural decisions
Seniority
Senior, hands-on IC