Senior Research Engineer, Model Evaluation
Core
Developing next-generation evaluation methods and scalable infrastructure to measure the performance and capabilities of frontier Large Language Models (LLMs).
Role type
Senior Research Engineer, Model Evaluation
Builds
Evaluation benchmarks, datasets, environments, and scalable analysis tools for LLMs.
Domain
Artificial Intelligence / Large Language Models (LLMs)
Deliverable
production ML models | research
Required skills
LLM evaluation method development, dataset creation, simulator/environment building, software engineering, performance analysis
Preferred skills
Publications at top-tier conferences, experience with popular benchmarks, deep LLM usage experience
Technologies
LLMs, evaluation frameworks, data pipelines
Responsibilities
Develop evaluation benchmarks, datasets, and environments for measuring bleeding-edge model capabilities; Conduct research to push the state-of-the-art in LLM evaluation methods, including training LLM judges; Build scalable tools for investigating and understanding evaluation results.
Seniority
Senior, hands-on IC