CareerPlanGet AI match score →

Site Reliability Engineer, Frontier Systems Infrastructure

San Francisco💼 Full-time🗓 2025-11-03 → 2026-07-31

Core

Operate and scale hyperscale supercomputers for frontier AI model training, managing the intersection of hardware and software at massive scale.

Role type

Senior Site Reliability Engineer (Hyperscale Infrastructure)

Builds

Next-generation compute clusters and software layers for large-scale model training

Domain

AI Research Infrastructure / Hyperscale Data Centers

Deliverable

infrastructure

Required skills

Kubernetes internals, container orchestration, Python/Go programming, Infrastructure-as-Code (Terraform/CloudFormation), bare-metal Linux, GPU hardware management, large-scale networking, firmware management, observability systems

Preferred skills

High-performance computing, cluster lifecycle automation, reducing operational latency

Technologies

Kubernetes, Terraform, CloudFormation, Linux, Python, Go

Responsibilities

Scale Kubernetes clusters to massive scale, automate bare-metal bring-up and provisioning, build software abstractions for training workloads, integrate networking and hardware health systems, develop monitoring and observability systems

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗