Principal Software Engineer, Inference
Core
Lead the model runtime architecture for HPE AI Essentials, an inference platform enabling enterprises to operate large language models on customer-owned hardware with a focus on sustained execution efficiency, low tail latency, and high GPU utilization.
Role type
Principal Software Engineer (LLM Inference Runtime)
Builds
HPE AI Essentials inference platform (model runtime, Kubernetes orchestration layer)
Domain
Enterprise AI / Large Language Model Inference / Cloud Infrastructure
Deliverable
production ML models
Required skills
LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM), continuous batching, KV cache management, quantization, speculative decoding, tensor/pipeline parallelism, Kubernetes operators/controllers, Go, Python, C++/CUDA profiling
Preferred skills
Upstream contributions to inference runtimes, disaggregated prefill/decode, RDMA/GPUDirect Storage, MIG/fractional GPU allocation, on-premises/air-gapped software delivery
Technologies
vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM, Kubernetes, NCCL, CUDA, Go, Python, C++, RDMA, InfiniBand, RoCE
Responsibilities
Define technical direction for LLM serving deployment including engine integration and distributed execution strategies; Partner with performance teams to optimize time-to-first-token and throughput; Evaluate and adopt emerging runtimes and serving techniques; Define orchestration layer for model admission, GPU scheduling, and autoscaling; Mentor engineers and lead architecture reviews.
Seniority
Principal, hands-on IC with mentorship