Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Core
Building and validating the end-to-end inference infrastructure for Claude on AWS, GCP, and Azure, ensuring correctness and performance for model launches and feature rollouts.
Role type
Staff/Senior Software Engineer, Cloud Inference Launch Engineering
Builds
Inference servers, load balancers, and CI/CD pipelines for LLM serving on major cloud platforms
Domain
Cloud Infrastructure, Large Language Model (LLM) Serving, Distributed Systems
Deliverable
production ML models
Required skills
High-performance distributed systems, automation/test infrastructure, cloud platform expertise (AWS/GCP/Azure), Kubernetes, Infrastructure as Code, container orchestration
Preferred skills
LLM inference optimization, capacity-constrained scheduling, multi-region deployments, request routing, Python, Rust
Technologies
AWS, GCP, Azure, Kubernetes, Infrastructure as Code
Responsibilities
Bring up inference for new model architectures and ship to cloud platforms; integrate new inference features (e.g., structured sampling, prompt caching); identify and fix cross-platform gaps in config, observability, and deployment; design and own CI/CD infrastructure with shadow traffic and correctness checks; drive down merge-to-production cycle time; analyze observability data to identify bottlenecks and drive remediation
Seniority
Staff/Senior, hands-on IC