Site Reliability Engineer (HPC)
Core
Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.
Role type
Site Reliability Engineer (HPC)
Builds
HPC clusters for ML model training and inference
Domain
High-Performance Computing / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes, Docker, container orchestration, CI/CD pipelines, public cloud platforms (Azure/AWS/GCP), infrastructure-as-code, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), Python, Go, Bash, distributed systems, networking, storage, GPU cluster management, workload schedulers, capacity planning
Preferred skills
ML training/inference pipelines, high-performance computing (HPC) and workload schedulers, background in capacity planning & cost optimization for GPU-heavy environments
Technologies
Kubernetes, Docker, Grafana, Datadog, OpenTelemetry, Azure, AWS, GCP
Responsibilities
Ensure uptime, resiliency, and fault tolerance of HPC clusters; Design and maintain monitoring, alerting, and logging systems; Build automation for deployments, incident response, scaling, and failover; Lead on-call rotations, troubleshoot production issues, and conduct blameless postmortems; Ensure data privacy, compliance, and secure operations; Partner with ML engineers and platform teams to improve developer experience
Seniority
Mid-level to Senior, hands-on IC