Engineering Manager, Inference Infrastructure
Core
Lead a team of ML platform and distributed-systems engineers to build the control plane for Anthropic's inference fleet, managing request routing, capacity allocation, and system performance.
Role type
Senior Engineering Manager, Inference Infrastructure
Builds
Control plane for inference fleet, load-balancing algorithms, and quantitative models for demand/capacity
Domain
AI/ML Infrastructure, Distributed Systems, Cloud Computing
Deliverable
production ML models | infrastructure
Required skills
Engineering management at scale, distributed systems architecture, load balancing, cluster orchestration, autoscaling, high-performance networking, incident response, quantitative modeling, team hiring and development
Preferred skills
LLM inference serving (KV caching, continuous batching), Kubernetes internals, multi-cloud workload management, heterogeneous accelerator fleet management, supercomputing/hyperscaler experience
Technologies
Kubernetes, Borg-style systems, cloud providers, heterogeneous accelerators
Responsibilities
Own technical roadmap for inference fleet coordination, partner with product/engineering teams for throughput/latency/cost wins, build habits of quantitative modeling, set technical strategy across hardware/clouds, run operational backbone (on-call, incident response), develop and hire strong teams, shape team structure during growth
Seniority
Senior, hands-on IC leader managing multiple teams