Staff Software Engineer, Inference Cloud
Core
Design and build the cloud layer for a globally distributed AI inference platform, ensuring high availability, low latency, and reliability for ultra-high-speed model serving.
Role type
Staff Software Engineer (Distributed Systems/Cloud Infrastructure)
Builds
Multi-region inference cloud platform, service discovery, request routing, load balancing, and traffic management systems.
Domain
AI Infrastructure / Cloud Systems / Distributed Computing
Deliverable
production ML models | infrastructure
Required skills
distributed systems architecture, cloud infrastructure, networking, compute orchestration, container platforms, Go, C++, Python, observability, incident response, SLO-driven operations, system design, code review, technical leadership
Preferred skills
ML inference infrastructure, model serving systems, GPU-accelerated workloads, TTFT optimization, tail-latency reduction
Technologies
Go, C++, Python, Kubernetes, gRPC, Prometheus, Grafana, AWS, GCP, Azure
Responsibilities
Shape technical direction for multi-region topology and system evolution; design critical components like service discovery and traffic management; architect active-active systems with rapid failover; define admission control and rate limiting mechanisms; write and review production code for critical paths; lead incident response and capacity planning; mentor senior engineers on design and standards.
Seniority
Staff, hands-on IC with strategic influence