Software Engineer, Safeguards Evals
Core
Build evaluation infrastructure for AI safety investigation agents to measure detection performance, robustness, and coverage against real-world misuse.
Role type
Senior IC machine-learning engineer (safety evals)
Builds
Evaluation harnesses, datasets, and pipelines for agentic investigation systems
Domain
AI safety, trust and safety, large language models
Deliverable
production ML models
Required skills
Python, data pipelines, LLMs and agentic systems, data analysis, research prototyping, experiment design
Preferred skills
agent evaluation frameworks, trust and safety, red teaming, synthetic data generation, distributed systems, prompt engineering
Technologies
Python, LLMs, RL environments
Responsibilities
Build and own evaluation harness for agentic investigation system; Construct high-quality eval datasets representing real-world misuse; Measure agent performance end-to-end; Analyze coverage to identify measurement gaps; Productionize research into regression and release pipelines; Build tooling for policy experts to author evaluations; Construct RL environments to improve safety investigation capabilities
Seniority
Senior, hands-on IC