Software Engineer Manager: Agentic Evaluation Platform
Core
Lead the engineering team building scalable backend services, compute orchestration, and evaluation harnesses to measure and validate Apple's Siri AI assistant capabilities.
Role type
Senior Engineering Manager (Agentic Evaluation Platform)
Builds
Scalable evaluation platform services, compute orchestration systems, and data workflows for AI agent validation
Domain
Consumer AI / Large Language Models / Test Infrastructure
Deliverable
production ML models
Required skills
Engineering management, technical leadership, distributed systems, cloud infrastructure, CI/CD, data pipelines, team building, ambiguity navigation
Preferred skills
LLM/Agentic evaluation, model experimentation frameworks, monitoring/logging/alerting, big data processing (Spark/Kafka/Flink), cross-functional leadership
Technologies
Python, Java, Go, Swift, Docker, Kubernetes, Apache Spark, Kafka, Flink, REST, gRPC
Responsibilities
Lead and grow a team of engineers; set technical direction and roadmap; own backend services and APIs for evaluation; manage compute orchestration across device fleets; define harness engineering contracts; steward interfaces between pipelines and consuming teams; own data workflows for results and reporting; deliver reliability dashboards; set standards for diagnosability; partner with QE and feature teams; recruit and mentor engineers
Seniority
Senior, hands-on IC with management responsibilities
