Evaluation & Insights Machine Learning Engineer
Core
Evaluate and improve AI systems (LLMs, multimodal models) by combining data science, model behavior analysis, and qualitative insights to ensure reliability, safety, and human alignment.
Role type
Senior IC machine learning engineer (evaluation & insights)
Builds
Scalable evaluation pipelines, automated annotation systems, and evaluation frameworks for Apple products
Domain
Artificial Intelligence / Machine Learning / Human-Centered AI
Deliverable
production ML models
Required skills
Python, PyTorch, JAX, Hugging Face, LLM evaluation frameworks, RAG architectures, fine-tuning, MLOps, distributed inference (Ray, vLLM), embedding-based clustering, prompt engineering
Preferred skills
Human factors, HCI, cognitive science methodologies
Technologies
Ray, vLLM, MLflow, Weights & Biases, LLM-as-a-judge, RLHF, DPO, G-Eval, DeepEval, vector databases
Responsibilities
Architect and execute comprehensive evaluation suites for LLMs and multimodal models; Develop deterministic, heuristic, and LLM-assisted evaluation frameworks; Translate qualitative failure modes into quantifiable loss patterns and guardrails; Partner with engineering to refine model behavior via telemetry; Apply advanced ML techniques to map error taxonomies and latent failure manifolds; Develop robust MLOps workflows for automated regression testing; Architect scalable distributed inference and processing pipelines; Define quantitative evaluation frameworks capturing human factors; Build automated evaluation pipelines utilizing LLMs; Collaborate cross-functionally to translate product requirements into evaluation infrastructure