CareerPlanSign in

Staff Software Engineer, Node Infra

New York City, NY💼 Full-time💰 $320,000–$320,000🗓 2026-04-30 → 2026-09-26

Core

Own the full lifecycle of accelerator capacity (ingestion, provisioning, scaling, health, diagnostics, repair) to power Anthropic's frontier AI research and scale Claude to millions of users.

Role type

Staff Infrastructure Engineer (Node Lifecycle & Cluster Orchestration)

Builds

Large-scale AI clusters across multiple clouds and accelerator families (GPUs, TPUs, Trainium) with automated health and repair systems.

Domain

AI Infrastructure / High-Performance Computing / Cloud Platforms

Deliverable

infrastructure

Required skills

distributed systems, reliability engineering, cloud platforms (AWS/GCP/Azure), systems programming (Rust/Go/Python), Infrastructure as Code (Terraform), machine learning accelerators, cross-team technical leadership, stakeholder alignment

Preferred skills

hyperscale compute management (10K+ nodes), Kubernetes internals (scheduler, autoscaler, Karpenter), cluster orchestration (Mesos, Borg), low-level systems (kernel, virtualization, device drivers), high-performance networking (EFA, RDMA, InfiniBand), production reliability for latency-sensitive systems, open-source contributions

Technologies

Kubernetes, Terraform, AWS, GCP, Azure, Rust, Go, Python, EFA, RDMA, InfiniBand, Mesos, Borg

Responsibilities

Own technical strategy and roadmap for node lifecycle management; drive cross-team initiatives to build and scale AI clusters; design systems for automatic hardware detection, isolation, and remediation; define infrastructure architecture; collaborate with cloud providers and internal teams on compute strategy; establish operational excellence practices; mentor engineers

Seniority

Staff, hands-on IC with strategic scope

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.