CareerPlanSign in

Software Engineer, Frontier Clusters Infrastructure

San Francisco💼 Full-time🗓 2024-11-07 → 2026-09-25

Core

Build, launch, and support hyperscale supercomputers for frontier AI model training by managing distributed systems and infrastructure at massive scale.

Role type

Senior IC distributed systems and infrastructure engineer

Builds

Next-generation compute clusters and software layers for large-scale AI model training

Domain

AI research infrastructure / Hyperscale data centers

Deliverable

infrastructure

Required skills

Kubernetes cluster scaling and lifecycle management, bare-metal provisioning and firmware upgrades, Python or Go programming, Infrastructure-as-Code (Terraform/CloudFormation), large-scale networking, GPU hardware management, observability system development

Preferred skills

Experience with high-performance computing (HPC), background in GPU workloads

Technologies

Kubernetes, Python, Go, Terraform, CloudFormation, Linux, GPU hardware

Responsibilities

Scale Kubernetes clusters to massive scale, automate bare-metal bring-up and cluster lifecycle management, build software abstractions to unify multiple clusters, improve operational metrics like cluster restart times, integrate networking and hardware health systems, develop monitoring and observability systems

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.