Staff Software Engineer, Node Infra
Core
Own the full lifecycle of accelerator capacity (ingestion, provisioning, scaling, health, diagnostics, repair) to power Anthropic's frontier AI research and scale Claude to millions of users.
Role type
Staff Infrastructure Engineer (Node Lifecycle & Cluster Orchestration)
Builds
Large-scale AI clusters across multiple clouds and accelerator families (GPUs, TPUs, Trainium) with automated health and repair systems.
Domain
AI Infrastructure / High-Performance Computing / Cloud Platforms
Deliverable
infrastructure
Required skills
distributed systems, reliability engineering, cloud platforms (AWS/GCP/Azure), systems programming (Rust/Go/Python), Infrastructure as Code (Terraform), machine learning accelerators, cross-team technical leadership, stakeholder alignment
Preferred skills
hyperscale compute management (10K+ nodes), Kubernetes internals (scheduler, autoscaler, Karpenter), cluster orchestration (Mesos, Borg), low-level systems (kernel, virtualization, device drivers), high-performance networking (EFA, RDMA, InfiniBand), production reliability for latency-sensitive systems, open-source contributions
Technologies
Kubernetes, Terraform, AWS, GCP, Azure, Rust, Go, Python, EFA, RDMA, InfiniBand, Mesos, Borg
Responsibilities
Own technical strategy and roadmap for node lifecycle management; drive cross-team initiatives to build and scale AI clusters; design systems for automatic hardware detection, isolation, and remediation; define infrastructure architecture; collaborate with cloud providers and internal teams on compute strategy; establish operational excellence practices; mentor engineers
Seniority
Staff, hands-on IC with strategic scope