Staff Infrastructure Engineer, Cluster Infrastructure
Core
Setting technical direction for agent-driven automation of compute cluster lifecycle management (provisioning, updates, decommissioning) across cloud providers and datacenters to support AI model training and scaling.
Role type
Staff Infrastructure Engineer (Cluster Infrastructure)
Builds
High-bandwidth, secure-by-default compute clusters interconnected across cloud providers and datacenters
Domain
Cloud Infrastructure / AI Compute Systems
Deliverable
infrastructure
Required skills
distributed systems, reliability engineering, cloud platforms, systems programming (Rust/Go/Python), Infrastructure as Code (Terraform)
Preferred skills
Kubernetes internals, cluster orchestration, cloud networking (VPC, BGP, Direct Connect), cluster security (RBAC, IAM, hardening), workflow orchestration, incident response
Technologies
Kubernetes, Terraform, AWS/GCP/Azure, Rust, Go, Python, Cilium, eBPF, Istio, Envoy, Temporal, Argo Workflows
Responsibilities
Own technical strategy and roadmap for cluster lifecycle management; Partner across teams to ingest new compute capacity; Align on physical build-out and cloud connectivity solutions; Collaborate with security to ensure secure-by-default provisioning; Define strategy for cluster scalability and fault tolerance; Work with cloud providers and internal teams on long-term compute strategy; Establish operational excellence practices including incident response and on-call health; Mentor engineers through technical coaching