AIML - Software Engineer - AI, Evaluation
Core
Design and build extensible frameworks, pipelines, and tools for the automatic evaluation of Apple AI products (Search, Siri, Apple Intelligence) using LLM-as-judge methodologies.
Role type
Senior IC AI software engineer (evaluation systems)
Builds
Evaluation frameworks, pipelines, and tools for AI model quality assessment
Domain
Consumer technology / Large Language Models / AI Evaluation
Deliverable
production ML models
Required skills
Python, system design, API design, CI/CD, testing strategies, code maintainability, debugging complex systems, AI-assisted development workflows
Preferred skills
LLM application development, MLOps, scalable evaluation tooling, statistical metrics interpretation (precision, recall, consistency), product requirement translation
Technologies
Python, LLM frameworks, MLOps tools
Responsibilities
Design and develop extensible frameworks for efficient AI model deployment and qualitative measurement; Build tools to support product launch decisions and cross-functional iteration; Implement LLM-as-judge systems for principled assessments of AI features