Senior Site Reliability Engineer - Fleet
Core
Build and operate monitoring, alerting, and automation for large-scale AI HPC clusters to ensure fabric, GPU, and job-level stability.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Production-grade AI cloud infrastructure (HPC clusters)
Domain
AI Cloud Infrastructure / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux distributed systems, InfiniBand/RoCE/CLOS fabrics, GPU-direct/NCCL, Python, Go, Ansible, Terraform, Prometheus, Grafana, Clickhouse, incident response, runbook creation
Preferred skills
PyTorch, TensorFlow, DeepSpeed, MLPerf, Docker, Kubernetes, NVIDIA hardware/firmware, data center power/thermal design, chaos engineering, SOC 2/ISO 27001 compliance
Technologies
InfiniBand, RoCE, CLOS, 100GbE, Ethernet, Ansible, Terraform, Prometheus, Grafana, Clickhouse, Python, Go, Docker, Kubernetes
Responsibilities
Build and operate monitoring and alerting for cluster health; Remotely deploy and configure large-scale HPC clusters; Automate cluster lifecycle (OS, firmware, drivers, networking); Create runbooks and automated remediations; Troubleshoot cluster issues across fabric, switching, and power; Participate in on-call rotations and lead incident response; Contribute to SOPs and provide requirements for operational efficiency
Seniority
Senior, hands-on IC