Member of Engineering (Evaluations)
Core
Design and implement infrastructure and tooling for AI evaluations and benchmarks to measure progress on real-world software development skills.
Role type
Senior IC machine-learning engineer (evaluations & infrastructure)
Builds
Evaluation frameworks, benchmarks, and tooling for LLMs
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
Large Language Models (LLM) expertise, strong programming skills across multiple languages, Linux proficiency, strong algorithmic skills, full software development life cycle experience, critical thinking
Preferred skills
Good taste, curious mindset, ability to question code quality policies
Technologies
Python, Linux
Responsibilities
Research and implementation of evaluations and benchmarks for base and instruction-following models, Collaborate with applied research and product teams to define meaningful metrics, Plan future steps and communicate clearly with peers
Seniority
Senior, hands-on IC