Manager; AI Evaluation Engineering
Core
Lead a team dedicated to evaluating and validating advanced generative AI solutions, including intelligent agents and digital assistants, ensuring robustness and reliability.
Role type
Manager, AI Evaluation Engineering
Builds
Enterprise-scale GenAI solutions across hybrid cloud and embedded/edge environments
Domain
Generative AI, AI Evaluation, Software Engineering
Deliverable
production ML models | product features
Required skills
Team leadership, AI evaluation methodologies, GenAI architecture design, Test automation, CI/CD, LLM/SLM expertise, RAG architecture, Vector databases, AI observability, Metric design, Human-in-the-loop validation
Preferred skills
Scaling engineering practices, Progressive deployment strategies, Prompt engineering, Agentic systems, Fine-tuning techniques (LoRA), MCP/A2A standards, Synthetic data creation, LLM-assisted evaluation
Technologies
Azure, AWS, GCP, Azure AI Foundry, SageMaker, Bedrock, Snowflake Cortex, LangChain, LangGraph, Langfuse, Arize, LangSmith, Humanloop, Ragas, DeepEval, Phoenix
Responsibilities
Provide technical direction to align the team with company goals, oversee individual and team performance, establish engineering best practices, ensure product quality and reliability
Seniority
Manager, hands-on leadership