Research Scientist / Engineer, Foundation Model Evaluation
Core
Design and implement evaluation systems, benchmarks, and metrics to measure frontier foundation model capabilities and drive model improvement for Apple products.
Role type
Senior IC research scientist / engineer (foundation model evaluation)
Builds
Evaluation benchmarks, metrics, test suites, and analysis frameworks for reasoning, code, knowledge, and agentic workflows
Domain
AI / Machine Learning / Natural Language Processing
Deliverable
production ML models
Required skills
Machine learning fundamentals, natural language processing, statistical analysis, Python, PyTorch or JAX, experimental design, benchmark design, model-based judging, tooling development
Preferred skills
PhD in CS/ML/NLP, experience evaluating large language models, human evaluation methodology, building reusable evaluation tooling, collaborating with model training teams
Technologies
Python, PyTorch, JAX
Responsibilities
Design and implement evaluation benchmarks and metrics; Develop product-aligned evaluation methods; Research and build evaluation tooling and analysis frameworks; Execute rigorous experiments and gap analysis; Collaborate with model training and product teams to inform development strategies
Seniority
Senior, hands-on IC