Senior Software Engineer, Inference
Core
Design and implement key components of the LLM inference runtime (engine integration, batching, KV cache management, distributed execution) to achieve low tail latency and high GPU utilization on customer-owned hardware.
Role type
Senior IC software engineer (LLM inference runtime)
Builds
Model runtime for HPE AI Essentials (inference platform for enterprises)
Domain
AI/ML infrastructure, cloud computing, high-performance computing
Deliverable
production ML models
Required skills
LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM), continuous batching, KV cache management, quantization, speculative decoding, tensor/pipeline parallelism, Kubernetes operators/controllers, Go, Python, C++/CUDA profiling
Preferred skills
Upstream contributions to inference runtimes, disaggregated prefill/decode, RDMA/GPUDirect Storage/InfiniBand, MIG/fractional GPU allocation, on-premises/air-gapped software delivery
Technologies
vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM, Kubernetes, NCCL, RDMA, InfiniBand, RoCE, CUDA, Nsight
Responsibilities
Design and implement LLM serving components (engine integration, batching, KV cache, quantized execution); Optimize time-to-first-token, inter-token latency, throughput, and tail latency; Build distributed execution capabilities (disaggregated prefill/decode, parallelism, cache offload); Evaluate emerging runtimes and quantization schemes; Contribute to orchestration layer (model admission, GPU scheduling, autoscaling); Triage and resolve customer issues; Mentor team members and lead engineering practices
Seniority
Senior, hands-on IC