CareerPlanGet AI match score →

Staff + Senior Software Engineer, Inference Deployment

San Francisco, CA💼 Full-time💰 $320,000–$320,000🗓 2026-06-29 → 2026-07-31

Core

Design and build deployment infrastructure that moves inference code from merge to production across GPU, TPU, and Trainium fleets, optimizing for resource constraints and minimizing disruption to live user traffic.

Role type

Staff + Senior Software Engineer (Inference Deployment)

Builds

Deployment orchestration systems, capacity-aware scheduling tools, observability dashboards, and self-service model onboarding pipelines.

Domain

AI Infrastructure / Cloud Systems / High-Performance Computing

Deliverable

production ML models

Required skills

Kubernetes deployment, container orchestration, complex state machine design, multi-stage pipeline architecture, backend services, CLI tools, web UIs

Preferred skills

Python, Rust, ML inference infrastructure, capacity planning, bin-packing, progressive delivery strategies, large-scale release engineering

Technologies

Kubernetes, GPU, TPU, Trainium

Responsibilities

Own unattended deployment orchestration across heterogeneous accelerator fleets; Improve capacity-aware scheduling to maximize throughput against constrained hardware budgets; Extend deployment observability with dashboards and tooling; Drive down cycle time from code merge to production via parallelized pipeline architectures; Optimize fleet rollout strategies for thousands of chips; Evolve self-service model onboarding; Partner with validation and autoscaling teams to integrate deployment automation.

Seniority

Staff + Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗