Software Engineer, Frontier Clusters Infrastructure
Core
Build, launch, and support hyperscale supercomputers for frontier AI model training by managing distributed systems and infrastructure at massive scale.
Role type
Senior IC distributed systems and infrastructure engineer
Builds
Next-generation compute clusters and software layers for large-scale AI model training
Domain
AI research infrastructure / Hyperscale data centers
Deliverable
infrastructure
Required skills
Kubernetes cluster scaling and lifecycle management, bare-metal provisioning and firmware upgrades, Python or Go programming, Infrastructure-as-Code (Terraform/CloudFormation), large-scale networking, GPU hardware management, observability system development
Preferred skills
Experience with high-performance computing (HPC), background in GPU workloads
Technologies
Kubernetes, Python, Go, Terraform, CloudFormation, Linux, GPU hardware
Responsibilities
Scale Kubernetes clusters to massive scale, automate bare-metal bring-up and cluster lifecycle management, build software abstractions to unify multiple clusters, improve operational metrics like cluster restart times, integrate networking and hardware health systems, develop monitoring and observability systems
Seniority
Senior, hands-on IC