Staff+ Software Engineer, ML Inference Path
Core
Design, build, and operate production infrastructure for ML-based safety systems (classifiers and defenses) powering Claude, serving thousands of models across all platforms.
Role type
Staff+ Software Engineer, ML Inference Path
Builds
Production infrastructure for safety classifiers, monitoring/observability tools, and deployment pipelines for AI safety systems.
Domain
AI Safety / Large-Scale Distributed Systems / Machine Learning
Deliverable
production ML models
Required skills
Python, PyTorch, TensorFlow, JAX, distributed systems, high-throughput low-latency system design, automated/self-service deployment pipelines, A/B testing frameworks, ML model evaluation infrastructure
Preferred skills
LLMs and transformer architectures, ML monitoring and alerting, data drift detection, trust & safety/fraud prevention/content moderation, privacy-preserving ML
Technologies
PyTorch, TensorFlow, JAX
Responsibilities
Design scalable ML infrastructure for real-time safety deployments; Build monitoring and observability tools for safety-critical applications; Collaborate with research teams to productionize safety techniques; Optimize inference latency and throughput for safety evaluations; Implement automated testing, deployment, and rollback systems; Partner with Security and Alignment teams to deliver safety infrastructure.
Seniority
Staff+, hands-on IC with strategic scope
