Staff + Sr. Software Engineer, Scaling
Core
Design, build, and maintain distributed systems that serve Claude to millions of users worldwide, focusing on intelligent request routing, load balancing, and fleet-wide orchestration across diverse AI accelerators.
Role type
Staff/Senior Software Engineer (Distributed Systems & Inference Infrastructure)
Builds
High-performance inference infrastructure, intelligent routing systems, and production-grade deployment pipelines for LLMs.
Domain
Artificial Intelligence / Large Language Model (LLM) Inference / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, high-performance computing, load balancing, request routing, cloud infrastructure management, Python, Rust
Preferred skills
LLM inference optimization, Kubernetes, multi-cloud orchestration, autoscaling strategies, observability analysis
Technologies
Kubernetes, AWS, GCP, Azure, Python, Rust
Responsibilities
Design intelligent request routing algorithms, develop autoscaling systems for compute fleets, build deployment pipelines for model releases, integrate new AI accelerator platforms, manage multi-region deployments, analyze observability data for performance tuning.
Seniority
Staff/Senior, hands-on IC