CareerPlanSign in

Staff+ Site Reliability Engineer, Safeguards ML Infra

San Francisco, CA💼 Full-time💰 $320,000–$320,000🗓 2026-09-04 → 2026-09-26

Core

Design, build, and operate production infrastructure for Claude's safety systems, ensuring safeguards are configured and deployed for every model launch.

Role type

Staff+ Site Reliability Engineer (ML Infra)

Builds

Production ML infrastructure for safety classifiers and model launch pipelines

Domain

AI Safety / Large Language Model (LLM) Infrastructure

Deliverable

production ML models

Required skills

Production change management at scale, deploy pipelines, config management systems, canary analysis, incident response, cloud platform operations (AWS, GCP), Python

Preferred skills

Rust, LLM inference systems, transformer architectures, reducing operational toil through automation, launch readiness review processes

Technologies

AWS, GCP, Python, Rust

Responsibilities

Stand up, configure, and verify safeguards for new model launches; deploy new safety classifiers via canary rollouts; detect and eliminate configuration drift across platforms; automate deployment pipelines and validation checks; maintain a safeguards registry with full provenance; participate in on-call rotations for service incidents and model provisioning.

Seniority

Staff+, hands-on IC with strategic impact

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.