Staff+ Software Engineer, Capacity Engineering
Core
Build production systems for data pipelines, observability tooling, and performance instrumentation to manage, plan, and maximize utilization of Anthropic's massive heterogeneous compute fleet (accelerators, CPUs, clouds).
Role type
Staff+ Software Engineer, Capacity Engineering
Builds
Data pipelines ingesting telemetry into BigQuery, real-time fleet health observability, capacity planning platforms, and utilization benchmarking infrastructure.
Domain
Cloud Infrastructure / Machine Learning Operations / FinOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, SQL, Kubernetes operations, Cloud provider operations (AWS/GCP/Azure), Observability (Prometheus/PromQL/Grafana), Data pipeline engineering, SLO management
Preferred skills
Capacity planning, Scheduling and packing efficiency, Multi-cloud data ingestion, Total cost of ownership forecasting, Accelerator infrastructure (GPU/TPU), Internal data product design, Exabyte-scale storage efficiency
Technologies
Python, SQL, BigQuery, Kubernetes, Prometheus, Grafana, AWS, GCP, Azure, DCGM
Responsibilities
Build planning and allocation stack for cross-region/cross-provider capacity; Drive efficiency programs (rightsizing, unused capacity recovery); Own attribution and forecasting by reconciling billing against telemetry; Build the underlying data platform with real-time SLOs; Operate Kubernetes-native systems at scale; Treat outputs as products with self-service access and documentation.