Benchmarking Project Lead, Siri Evaluation
Core
Design and execute human-in-the-loop evaluation processes to benchmark the accuracy of next-generation Siri AI across features, locales, and platforms.
Role type
Senior Manager, Evaluation & Data Science
Builds
Reliable evaluation systems and tooling ecosystems for managing rich datasets
Domain
Consumer AI / Human-in-the-loop evaluation
Deliverable
production ML models
Required skills
Agentic coding, metrics analysis, crowd science, data collection, annotation analysis, statistics, budget management, cross-functional collaboration
Preferred skills
Computational linguistics, language quality assessment, Python, data pipeline engineering
Responsibilities
Design efficient data collection processes using humans in the loop; Lead annotation efforts for various languages and devices; Plan and manage annotator resourcing budgets; Track and improve quality of human judgements via training and review mechanisms; Collaborate with engineering teams to build tooling ecosystems.
Seniority
Senior, hands-on IC with management responsibilities