Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations
Core
Lead AI evaluation practices to ensure reliability, safety, and accuracy of autonomous agentic workflows for enterprise customers.
Role type
Senior Engineering Manager, Agentic & GenAI Benchmarking and Evaluations
Builds
Automated testing/evaluation harnesses, enterprise AI benchmarks, and validation frameworks for multi-agent orchestration and RAG pipelines.
Domain
Enterprise AI, Agentic Workflows, Generative AI, MLOps
Deliverable
production ML models | product features
Required skills
Multi-agent orchestration, probabilistic software architectures, rigorous AI metrics (ROUGE, BLEU, G-Eval), Python, SQL, MLOps platforms, knowledge graph architectures, executive risk assessment
Preferred skills
Master's or Ph.D. in CS/Data Science/ML, experience with frontier LLMs (OpenAI, Anthropic, Google)
Technologies
Python, SQL, Pandas, NumPy, MLOps tracking platforms
Responsibilities
Design and scale automated testing/evaluation harnesses; Create standard, repeatable evaluation frameworks for complex business workflows; Validate RAG pipelines, hybrid search, and semantic re-ranking systems; Benchmark frontier LLMs for execution capabilities and costs; Recruit and mentor an AI-native engineering team; Translate technical evaluation data into executive-level risk assessments and strategic recommendations
Seniority
Senior, hands-on IC with management responsibilities