Staff Infrastructure Engineer, Cluster Infrastructure
Core
Set technical direction for agent-driven automation of compute cluster lifecycle management (provisioning, updates, decommissioning) across major cloud providers and datacenters to support training new models and scaling Claude. (via careerplan.io/jobs/5211297008-staff-infrastructure-engineer-cluster-infrastructure-at-anthropic)
Role type
Staff Infrastructure Engineer (Cluster Infrastructure)
Builds
High-bandwidth, secure-by-default compute clusters interconnected across clouds and datacenters
Domain
Cloud Infrastructure / Distributed Systems / AI Compute
Deliverable
production ML models
Required skills
Distributed systems expertise, Cloud platforms (AWS/GCP/Azure), Systems languages (Rust/Go/Python), Infrastructure as Code (Terraform), Kubernetes internals, Cluster orchestration, Cloud networking (VPC, BGP, Direct Connect), Cluster security (RBAC, IAM, hardening), Workflow orchestration
Preferred skills
Hyperscale compute operations (100+ clusters, 10K+ nodes), eBPF, Service mesh (Istio/Envoy), Supply-chain security, Rapid systems design adaptation
Technologies
Kubernetes, Terraform, AWS, GCP, Azure, Rust, Go, Python, Cilium, eBPF, Istio, Envoy, Temporal, Argo Workflows
Responsibilities
Own technical strategy and roadmap for cluster lifecycle management, Partner across teams to ingest new compute capacity, Align on physical build-out and cloud connectivity solutions, Collaborate with security to ensure secure-by-default provisioning, Define strategy on cluster scalability and fault tolerance, Work with cloud providers and internal teams on long-term compute strategy, Establish operational-excellence practices (incident response, postmortem culture), Mentor and coach engineers
Seniority
Staff, technical direction & mentorship