Staff Software Engineer, AI Reliability Engineering
Core
Building reliable, interpretable, and steerable AI systems (Claude) by improving reliability across serving paths, infrastructure, and accelerators.
Role type
Staff Software Engineer, AI Reliability Engineering
Builds
High-availability serving infrastructure, monitoring/observability systems, and incident response capabilities for large language model serving.
Domain
Artificial Intelligence / Large Language Model Serving / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
distributed systems, infrastructure, reliability engineering, incident response, high-availability system design, cross-team collaboration, holistic system architecture
Preferred skills
SRE or Production Engineer experience, large-scale model serving or training infrastructure operation, ML hardware accelerator expertise (GPUs, TPUs, Trainium), ML-specific networking optimizations (RDMA, InfiniBand), AI-specific observability tools, chaos engineering, open-source infrastructure contributions
Technologies
RDMA, InfiniBand, GPUs, TPUs, Trainium
Responsibilities
Develop Service Level Objectives for LLM serving systems; Design and implement monitoring and observability systems across the token path; Assist in designing high-availability serving infrastructure across regions and cloud providers; Lead incident response for critical AI services; Support reliability of safeguard model serving.
Seniority
Staff, hands-on IC with strategic cross-cutting impact