Staff Machine Learning Platform Engineer, AI Evaluation
Core
Lead architectural design and development of high-availability AI evaluation systems to enable self-service evaluation at scale for Apple's generative AI and agent systems.
Role type
Staff Machine Learning Platform Engineer (AI Evaluation)
Builds
APIs, SDKs, orchestration services, and reusable abstractions for operationalizing ML research into production services.
Domain
Generative AI, Agent Systems, ML Infrastructure
Deliverable
production ML models
Required skills
Python (FastAPI, Pydantic), job orchestration (Temporal.io), distributed compute (Ray), CI/CD, containerization (Docker/K8s), monitoring, system design, strategic decision-making, research code assessment, AI/agent evaluation logic.
Preferred skills
Go, Rust, startup/early-stage experience, LLM token economics, rate limiting, cost management at scale.
Responsibilities
Own technical direction for the evaluation platform; partner with researchers to productionize ML code; balance competing priorities from engineering teams and leadership; define org-level evaluation strategy; own developer experience for evaluation patterns; establish operational rigor for testing and reliability.
Seniority
Staff, hands-on IC with strategic scope