AIML - Manager, Applied AI Science - GenAI Model Autograding, Evaluation
Core
Lead a team building state-of-the-art autograder systems to evaluate the quality of internal GenAI products (Search, Siri, Apple Intelligence) using advanced AI/ML techniques.
Role type
Manager, Applied AI Science (GenAI Model Autograding, Evaluation)
Builds
Reliable, scalable LLM/VLM-based evaluation systems and autograder toolings for product-quality decisions.
Domain
Generative AI, Large Language Models (LLM), Vision-Language Models (VLM), Model Evaluation
Deliverable
production ML models
Required skills
Team leadership (hiring, development, performance management), Technical vision & roadmap setting, LLM/VLM-based evaluation system design, Rubric design & evaluation set development, Model alignment & calibration, Cross-functional stakeholder alignment
Preferred skills
Leading complex ambiguous AI/ML projects from definition to delivery, Scaling AI/ML solutions into reusable platforms/frameworks, Developing senior ML engineers/technical leads
Technologies
LLMs, VLMs, Autograders
Responsibilities
Lead and grow a high-performing team of applied AI scientists and ML engineers; Set technical vision and roadmap for autograder development; Deliver reliable, scalable evaluation systems; Navigate cross-functional dependencies; Standardize and automate end-to-end evaluation workflows
Seniority
Manager, hands-on technical leadership

