CareerPlanSign in

Principal Software Engineer, Inference

Spring, Texas, United States of America💼 Full-time💰 $152,000–$152,000🗓 2026-09-16 → 2026-09-25

Core

Lead the model runtime architecture for HPE AI Essentials, an inference platform enabling enterprises to operate large language models on customer-owned hardware with a focus on sustained execution efficiency, low tail latency, and high GPU utilization.

Role type

Principal Software Engineer (LLM Inference Runtime)

Builds

HPE AI Essentials inference platform (model runtime, Kubernetes orchestration layer)

Domain

Enterprise AI / Large Language Model Inference / Cloud Infrastructure

Deliverable

production ML models

Required skills

LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM), continuous batching, KV cache management, quantization, speculative decoding, tensor/pipeline parallelism, Kubernetes operators/controllers, Go, Python, C++/CUDA profiling

Preferred skills

Upstream contributions to inference runtimes, disaggregated prefill/decode, RDMA/GPUDirect Storage, MIG/fractional GPU allocation, on-premises/air-gapped software delivery

Technologies

vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM, Kubernetes, NCCL, CUDA, Go, Python, C++, RDMA, InfiniBand, RoCE

Responsibilities

Define technical direction for LLM serving deployment including engine integration and distributed execution strategies; Partner with performance teams to optimize time-to-first-token and throughput; Evaluate and adopt emerging runtimes and serving techniques; Define orchestration layer for model admission, GPU scheduling, and autoscaling; Mentor engineers and lead architecture reviews.

Seniority

Principal, hands-on IC with mentorship

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.