Staff Software Engineer, Inference
Core
Design and operate a Kubernetes-native inference platform for massive-scale GPU workloads, focusing on low-latency optimization, cost efficiency, and reliability for AI/ML, VFX, and real-time inference.
Role type
Staff Software Engineer (IC5), technical leader
Builds
Core cloud platform for GPU inference (request routing, adaptive scheduling, GPU resource management)
Domain
Cloud Infrastructure / AI/ML Inference
Deliverable
production ML models | infrastructure
Required skills
Go, Python, C++, Kubernetes (orchestration, scheduling), distributed systems design, networked systems, performance optimization, inference system optimization (batching, caching, memory, mixed precision, streaming), SLI/SLO ownership, capacity planning, autoscaling, mentoring
Preferred skills
Open-source contributions to inference frameworks (vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe), GPU systems engineering (CUDA, NCCL, RDMA, NUMA), hyperscale cloud experience
Responsibilities
Define and lead cross-cutting design initiatives for inference platform architecture, implement advanced inference optimizations (speculative decoding, KV-cache reuse), establish performance benchmarking frameworks, guide cross-functional alignment, improve tail latency and platform reliability, mentor senior and mid-level engineers
Seniority
Staff, technical leader across multiple teams