Senior Software Engineer, Inference
Core
Lead design and optimization of a Kubernetes-native inference platform to serve AI labs, startups, and enterprises with high-performance, low-latency model serving.
Role type
Senior IC Inference Software Engineer
Builds
High-scale, low-latency AI inference services and orchestration systems
Domain
Cloud Infrastructure / AI Inference / Distributed Systems
Deliverable
production ML models
Required skills
Python, Go, CUDA kernel development, Kubernetes, networked systems, performance optimization, tail latency reduction, batching, caching, mixed precision, streaming token delivery
Preferred skills
vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe, NCCL, RDMA, GPU interconnect topologies, multi-team leadership
Technologies
Kubernetes, Prometheus, Grafana, OpenTelemetry, CUDA, NCCL, RDMA
Responsibilities
Lead design reviews and drive architecture; define and own SLIs/SLOs; implement advanced optimizations (micro-batch schedulers, speculative decoding, KV-cache reuse); strengthen incident posture (capacity planning, autoscaling, graceful degradation); mentor engineers and review cross-team designs; own an area spanning multiple services and teams
Seniority
Senior, hands-on IC