ML Infrastructure Engineer, Safeguards
Core
Build and scale critical ML infrastructure to power AI safety systems, including real-time and batch classifiers for model evaluation.
Role type
Senior ML Infrastructure Engineer (AI Safety)
Builds
Scalable platforms and tools for safety evaluations, monitoring, and observability for large-scale AI models.
Domain
AI Safety / Large-scale Distributed Systems
Deliverable
production ML models
Required skills
Python, PyTorch/TensorFlow/JAX, Cloud platforms (AWS/GCP), Kubernetes, Distributed systems, Data engineering (Spark/Airflow/streaming), Automated testing/deployment
Preferred skills
LLMs/Transformers, A/B testing frameworks, ML monitoring/alerting, Human-in-the-loop workflows, Trust & safety domains, Privacy-preserving ML, Open-source ML infrastructure
Technologies
Python, PyTorch, TensorFlow, JAX, AWS, GCP, Kubernetes, Spark, Airflow
Responsibilities
Design scalable ML infrastructure for real-time/batch safety evaluations; Build monitoring/observability tools for model performance and data quality; Collaborate with research to productionize safety techniques; Optimize inference latency/throughput; Implement automated testing/deployment/rollback systems; Partner with Safety/Security/Alignment teams; Develop internal tools for safety research.
Seniority
Senior, hands-on IC