CareerPlanSign in

Senior Site Reliability Engineer - Fleet

San Francisco Office (Fremont St)🌐 Remote💼 Full-time🗓 2026-09-01 → 2026-09-26

Core

Build and operate monitoring, alerting, and automation for large-scale AI HPC clusters to ensure fabric, GPU, and job-level stability.

Role type

Senior Site Reliability Engineer (AI Infrastructure)

Builds

Production-grade AI cloud infrastructure (HPC clusters)

Domain

AI Cloud Infrastructure / High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Linux distributed systems, InfiniBand/RoCE/CLOS fabrics, GPU-direct/NCCL, Python, Go, Ansible, Terraform, Prometheus, Grafana, Clickhouse, incident response, runbook creation

Preferred skills

PyTorch, TensorFlow, DeepSpeed, MLPerf, Docker, Kubernetes, NVIDIA hardware/firmware, data center power/thermal design, chaos engineering, SOC 2/ISO 27001 compliance

Technologies

InfiniBand, RoCE, CLOS, 100GbE, Ethernet, Ansible, Terraform, Prometheus, Grafana, Clickhouse, Python, Go, Docker, Kubernetes

Responsibilities

Build and operate monitoring and alerting for cluster health; Remotely deploy and configure large-scale HPC clusters; Automate cluster lifecycle (OS, firmware, drivers, networking); Create runbooks and automated remediations; Troubleshoot cluster issues across fabric, switching, and power; Participate in on-call rotations and lead incident response; Contribute to SOPs and provide requirements for operational efficiency

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.