Senior Backend Engineer, Inference Platform
Core
Building the core inference backbone and optimizing global request routing, load balancing, and resource allocation for multi-tenant serverless workloads serving frontier generative AI models.
Role type
Senior Backend Engineer, Inference Platform
Builds
Multi-tenant serverless inference platform powering LLMs, multimodal, image, audio, video, and speech models at scale.
Domain
Generative AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed systems design, API microservices, low-level OS concepts (multi-threading, memory management, networking, storage), Rust/Go/Python/TypeScript, system profiling, auto-scaling, prefix caching optimization, GPU software stacks (CUDA, Triton, NCCL), HPC technologies (InfiniBand, NVLink, MPI), Kubernetes/container orchestration, open source inference tools (SGLang, vLLM, NVIDIA Dynamo)
Preferred skills
Knowledge of modern LLMs and generative model serving, experience with open source inference ecosystem
Technologies
H100, H200, GB200 GPUs, SGLang, vLLM, NVIDIA Dynamo, Kubernetes, CUDA, Triton, NCCL, InfiniBand, NVLink, MPI
Responsibilities
Build and optimize global and local request routing with low-latency load balancing across data centers; Develop auto-scaling systems to dynamically allocate resources and meet strict SLOs; Design systems for multi-tenant traffic shaping, rate limiting, and resource allocation; Engineer trade-offs between latency and throughput for diverse workloads; Optimize prefix caching to reduce model compute and speed up responses; Collaborate with ML researchers to bring new model architectures into production; Continuously profile and analyze system-level performance to identify bottlenecks and implement optimizations
Seniority
Senior, hands-on IC