Staff + Senior Software Engineer, Inference Deployment
Core
Design and build deployment infrastructure that moves inference code from merge to production across GPU, TPU, and Trainium fleets, optimizing for resource constraints and minimizing disruption to live user traffic.
Role type
Staff + Senior Software Engineer (Inference Deployment)
Builds
Deployment orchestration systems, capacity-aware scheduling tools, observability dashboards, and self-service model onboarding pipelines.
Domain
AI Infrastructure / Cloud Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Kubernetes deployment, container orchestration, complex state machine design, multi-stage pipeline architecture, backend services, CLI tools, web UIs
Preferred skills
Python, Rust, ML inference infrastructure, capacity planning, bin-packing, progressive delivery strategies, large-scale release engineering
Technologies
Kubernetes, GPU, TPU, Trainium
Responsibilities
Own unattended deployment orchestration across heterogeneous accelerator fleets; Improve capacity-aware scheduling to maximize throughput against constrained hardware budgets; Extend deployment observability with dashboards and tooling; Drive down cycle time from code merge to production via parallelized pipeline architectures; Optimize fleet rollout strategies for thousands of chips; Evolve self-service model onboarding; Partner with validation and autoscaling teams to integrate deployment automation.
Seniority
Staff + Senior, hands-on IC