Member of Technical Staff - Evaluations
Core
Conduct comparative analysis and build evaluation systems to measure LLM capabilities, reasoning, and alignment.
Role type
Senior IC machine-learning engineer (evaluation)
Builds
Generalizable evaluation frameworks and feedback loops for model improvement
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
statistical analysis, experimental design, LLM evaluation methodologies, synthetic eval design, human feedback integration
Preferred skills
agentic task evaluation, real-world interaction data analysis
Technologies
LLMs, synthetic data, human feedback systems
Responsibilities
Conduct critical comparative analysis to advance understanding of model capabilities; Build and refine evaluation systems creating feedback loops between data, evals, and model behavior; Develop generalizable evaluation frameworks for reasoning, alignment, and usefulness; Collaborate with pre-training and post-training teams to translate insights into model improvements; Push boundaries of measurable metrics from synthetic evals to real-world interaction data
Seniority
Senior, hands-on IC