CareerPlanGet AI match score →

Staff Software Engineer, Inference Cloud

Headquarters/Sunnyvale Office💼 Full-time🗓 2024-07-12 → 2026-07-30

Core

Design and build the cloud layer for a globally distributed AI inference platform, ensuring high availability, low latency, and reliability for ultra-high-speed model serving.

Role type

Staff Software Engineer (Distributed Systems/Cloud Infrastructure)

Builds

Multi-region inference cloud platform, service discovery, request routing, load balancing, and traffic management systems.

Domain

AI Infrastructure / Cloud Systems / Distributed Computing

Deliverable

production ML models | infrastructure

Required skills

distributed systems architecture, cloud infrastructure, networking, compute orchestration, container platforms, Go, C++, Python, observability, incident response, SLO-driven operations, system design, code review, technical leadership

Preferred skills

ML inference infrastructure, model serving systems, GPU-accelerated workloads, TTFT optimization, tail-latency reduction

Technologies

Go, C++, Python, Kubernetes, gRPC, Prometheus, Grafana, AWS, GCP, Azure

Responsibilities

Shape technical direction for multi-region topology and system evolution; design critical components like service discovery and traffic management; architect active-active systems with rapid failover; define admission control and rate limiting mechanisms; write and review production code for critical paths; lead incident response and capacity planning; mentor senior engineers on design and standards.

Seniority

Staff, hands-on IC with strategic influence

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗