Research Engineer (Evals)
Core
Build and maintain internal benchmark suites for single/multi-turn content and agentic guardrails, and study agent behaviors in the wild.
Role type
Research Engineer (Evals)
Builds
Internal benchmark suites, evals for flagship models, and tooling for studying agent reliability.
Domain
AI Safety, LLM Evaluation, Agentic Systems
Deliverable
production ML models | product features
Required skills
LLM benchmark construction, synthetic data generation for post-training, Python production code development, efficient LLM inference orchestration, frontier model usage, coding agents
Preferred skills
Automated red-teaming, experience with agentic scaffolds, reward-model/safety benchmark knowledge, published papers in evals/safety-evaluation
Technologies
Python, LLMs, coding agents
Responsibilities
Own and maintain internal benchmark suite, build benchmarks distinguishing specific model capabilities, work with product team on flagship model evals, build benchmarks for new research features, adapt evals to new verticals, study and quantify realistic agentic/LLM failure modes
Seniority
Mid-Senior, hands-on IC
