CareerPlanGet AI match score →

Principal LLM Inference Engineer

Santa Clara💼 Full-time🗓 2026-06-30 → 2026-07-31

Core

End-to-end inference engineer optimizing LLM inference on heterogeneous hardware (silicon, CPUs, GPUs) from kernel-level to distributed orchestration.

Role type

Principal LLM Inference Engineer

Builds

Inference runtimes, serving frameworks, custom kernels, and proof-of-concept systems for frontier AI models.

Domain

Generative AI / Heterogeneous Compute / Systems Engineering

Deliverable

production ML models

Required skills

Python, C/C++, LLM inference optimization, CUDA/Triton, vLLM, SGLang, TensorRT-LLM, quantization, distributed inference

Preferred skills

Heterogeneous compute deployments, custom silicon/ASIC inference, speculative decoding, JAX Scaling Book knowledge

Technologies

vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX

Responsibilities

Prototype emerging LLM inference use cases for heterogeneous hardware, develop and tune custom kernels and operator-level optimizations, drive quantization and sparsity strategies, build and maintain inference runtimes and serving frameworks, contribute to distributed inference systems, partner with hardware architects and product teams.

Seniority

Principal, hands-on IC with strategy & mentorship

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗