Research Engineer, Knowledge Foundations
Core
Design and run experiments to improve Claude's ability to search, retrieve, and reason over knowledge-intensive tasks at scale.
Role type
Senior Research Engineer (LLM training & evaluation)
Builds
Training environments, data pipelines, evaluation suites, and infrastructure for LLM post-training
Domain
Artificial Intelligence / Large Language Models / Knowledge Work
Deliverable
production ML models
Required skills
Python engineering, ML experiment design, distributed systems, data pipeline development, model training, evaluation design, observability tooling
Preferred skills
RL training on LLMs, LLM evaluation in open-ended domains, experience in frontier AI labs, published research on LLMs/RL/retrieval, distributed training systems expertise
Technologies
Python, LLMs, RL, distributed training systems
Responsibilities
Design and iterate training environments and data pipelines; Run end-to-end ML experiments from hypothesis to analysis; Develop evaluations for search, retrieval, and reasoning quality; Identify failure modes and translate them into training signals; Collaborate with researchers to align priorities; Build shared infrastructure and tooling; Own evaluation tools and model release processes; Build observability dashboards and operational tooling
Seniority
Senior, hands-on IC