CareerPlanSign in

Customer Reliability Engineer

San Francisco, CA💼 Full-time🗓 2026-07-17 → 2026-09-26

Core

Own reliability for named customer workloads, debug distributed systems across the full stack, and manage customer-facing incident communication for large-scale AI compute clusters.

Role type

Senior Customer Reliability Engineer (AI Infrastructure)

Builds

Reliable, high-scale AI training clusters and compute infrastructure for frontier AI labs and cloud providers.

Domain

AI Infrastructure / High-Performance Computing (HPC) / Data Center Operations

Deliverable

production ML models | infrastructure

Required skills

Debugging distributed systems, incident management, cross-stack troubleshooting, customer communication, pushing engineering fixes

Preferred skills

GPU training workloads, InfiniBand, RoCE, Slurm, Kubernetes, NCCL debugging

Technologies

InfiniBand, RoCE, Slurm, Kubernetes, NCCL

Responsibilities

Own reliability for named customer workloads including clusters and SLAs, debug distributed systems across hardware, fabric, and scheduler layers, run customer-facing incident communication with technical depth, turn recurring customer pain into engineering fixes

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.