Staff+ Site Reliability Engineer, Safeguards ML Infra
Core
Design, build, and operate production infrastructure for Claude's safety systems, ensuring safeguards are configured and deployed for every model launch.
Role type
Staff+ Site Reliability Engineer (ML Infra)
Builds
Production ML infrastructure for safety classifiers and model launch pipelines
Domain
AI Safety / Large Language Model (LLM) Infrastructure
Deliverable
production ML models
Required skills
Production change management at scale, deploy pipelines, config management systems, canary analysis, incident response, cloud platform operations (AWS, GCP), Python
Preferred skills
Rust, LLM inference systems, transformer architectures, reducing operational toil through automation, launch readiness review processes
Technologies
AWS, GCP, Python, Rust
Responsibilities
Stand up, configure, and verify safeguards for new model launches; deploy new safety classifiers via canary rollouts; detect and eliminate configuration drift across platforms; automate deployment pipelines and validation checks; maintain a safeguards registry with full provenance; participate in on-call rotations for service incidents and model provisioning.
Seniority
Staff+, hands-on IC with strategic impact
