Site Reliability Engineer, Frontier Systems Infrastructure
Core
Operate and scale hyperscale supercomputers for frontier AI model training, managing the intersection of hardware and software at massive scale.
Role type
Senior Site Reliability Engineer (Hyperscale Infrastructure)
Builds
Next-generation compute clusters and software layers for large-scale model training
Domain
AI Research Infrastructure / Hyperscale Data Centers
Deliverable
infrastructure
Required skills
Kubernetes internals, container orchestration, Python/Go programming, Infrastructure-as-Code (Terraform/CloudFormation), bare-metal Linux, GPU hardware management, large-scale networking, firmware management, observability systems
Preferred skills
High-performance computing, cluster lifecycle automation, reducing operational latency
Technologies
Kubernetes, Terraform, CloudFormation, Linux, Python, Go
Responsibilities
Scale Kubernetes clusters to massive scale, automate bare-metal bring-up and provisioning, build software abstractions for training workloads, integrate networking and hardware health systems, develop monitoring and observability systems
Seniority
Senior, hands-on IC