Customer Reliability Engineer
Core
Own reliability for named customer workloads, debug distributed systems across the full stack, and manage customer-facing incident communication for large-scale AI compute clusters.
Role type
Senior Customer Reliability Engineer (AI Infrastructure)
Builds
Reliable, high-scale AI training clusters and compute infrastructure for frontier AI labs and cloud providers.
Domain
AI Infrastructure / High-Performance Computing (HPC) / Data Center Operations
Deliverable
production ML models | infrastructure
Required skills
Debugging distributed systems, incident management, cross-stack troubleshooting, customer communication, pushing engineering fixes
Preferred skills
GPU training workloads, InfiniBand, RoCE, Slurm, Kubernetes, NCCL debugging
Technologies
InfiniBand, RoCE, Slurm, Kubernetes, NCCL
Responsibilities
Own reliability for named customer workloads including clusters and SLAs, debug distributed systems across hardware, fabric, and scheduler layers, run customer-facing incident communication with technical depth, turn recurring customer pain into engineering fixes
Seniority
Senior, hands-on IC
