AI QA & Evaluation Engineer
Core
Design and implement test strategies, evaluation rubrics, and automation frameworks to validate the accuracy, reliability, and ethical use of GenAI models and infrastructure.
Role type
Senior IC AI QA & Evaluation Engineer
Builds
Robust testing frameworks, evaluation tasks, and CI/CD pipelines for GenAI solutions
Domain
Generative AI, Search AI, Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, TypeScript, LLM evaluation frameworks, Rubric-based evaluation design, RAG architecture, CI/CD automation, Cloud platforms (Azure/GCP/AWS), Data validation
Preferred skills
LangSmith, Confident AI, Azure OpenAI, Vertex AI, Terraform, GitHub
Technologies
LangSmith, Confident AI, Azure OpenAI, Vertex AI, ChatGPT Enterprise, Terraform, GitHub, Azure, GCP, AWS
Responsibilities
Design comprehensive test strategies for AI/ML systems including accuracy and bias testing; Create self-contained evaluation tasks and grading rubrics; Automate validation suites for agentic systems and CI/CD pipelines; Validate data sources for AI/ML model consumption; Document AI agent behaviors and model performance reports; Ensure AI data sources meet governance and compliance standards
Seniority
Senior, hands-on IC