Staff Software Engineer, Inference
Core
Building and maintaining critical inference systems to serve Claude to millions of users, focusing on compute efficiency and enabling breakthrough AI research.
Role type
Staff Software Engineer (Inference Infrastructure)
Builds
Compute-agnostic inference deployments, intelligent request routing, fleet-wide orchestration across diverse AI accelerators, and production-grade deployment pipelines.
Domain
Artificial Intelligence / Large Language Model (LLM) Inference / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems, Large-scale service orchestration, Intelligent request routing, Performance optimization, Multi-accelerator deployments, Cloud infrastructure (AWS, GCP), Python or Rust
Preferred skills
LLM inference optimization, Batching strategies, Caching strategies, Load balancing, Traffic management systems, Kubernetes
Responsibilities
Designing intelligent routing algorithms, Autoscaling compute fleets, Building deployment pipelines, Integrating new AI accelerator platforms, Contributing to inference features, Supporting new model architectures, Analyzing observability data, Managing multi-region deployments
Seniority
Staff, hands-on IC with strategic impact