Principal LLM Inference Engineer
Core
End-to-end inference engineer optimizing LLM inference on heterogeneous hardware (silicon, CPUs, GPUs) from kernel-level to distributed orchestration.
Role type
Principal LLM Inference Engineer
Builds
Inference runtimes, serving frameworks, custom kernels, and proof-of-concept systems for frontier AI models.
Domain
Generative AI / Heterogeneous Compute / Systems Engineering
Deliverable
production ML models
Required skills
Python, C/C++, LLM inference optimization, CUDA/Triton, vLLM, SGLang, TensorRT-LLM, quantization, distributed inference
Preferred skills
Heterogeneous compute deployments, custom silicon/ASIC inference, speculative decoding, JAX Scaling Book knowledge
Technologies
vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
Responsibilities
Prototype emerging LLM inference use cases for heterogeneous hardware, develop and tune custom kernels and operator-level optimizations, drive quantization and sparsity strategies, build and maintain inference runtimes and serving frameworks, contribute to distributed inference systems, partner with hardware architects and product teams.
Seniority
Principal, hands-on IC with strategy & mentorship