CareerPlanSign in

Staff + Sr. Software Engineer, Scaling

San Francisco, CA💼 Full-time💰 $320,000–$320,000🗓 2026-08-24 → 2026-09-26

Core

Design, build, and maintain distributed systems that serve Claude to millions of users worldwide, focusing on intelligent request routing, load balancing, and fleet-wide orchestration across diverse AI accelerators.

Role type

Staff/Senior Software Engineer (Distributed Systems & Inference Infrastructure)

Builds

High-performance inference infrastructure, intelligent routing systems, and production-grade deployment pipelines for LLMs.

Domain

Artificial Intelligence / Large Language Model (LLM) Inference / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Distributed systems design, high-performance computing, load balancing, request routing, cloud infrastructure management, Python, Rust

Preferred skills

LLM inference optimization, Kubernetes, multi-cloud orchestration, autoscaling strategies, observability analysis

Technologies

Kubernetes, AWS, GCP, Azure, Python, Rust

Responsibilities

Design intelligent request routing algorithms, develop autoscaling systems for compute fleets, build deployment pipelines for model releases, integrate new AI accelerator platforms, manage multi-region deployments, analyze observability data for performance tuning.

Seniority

Staff/Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.