Software Engineer- Inference Performance
Core
Build and optimize the inference engine and runtime to make demanding AI workloads run faster and more efficiently for customers.
Role type
Senior IC inference performance engineer (LLM)
Builds
High-performance inference stack, runtime internals, and scheduling systems for LLMs
Domain
AI/ML infrastructure, Large Language Models (LLM), GPU computing
Deliverable
production ML models
Required skills
C++, Python, LLM optimization techniques, PyTorch, TensorRT, GPU architecture, speculative decoding, quantization, KV-cache management, cross-layer profiling
Preferred skills
CUDA/Triton kernel development, upstream open-source contributions (vLLM, SGLang), large-scale distributed serving, FP8/FP4 quantization
Technologies
vLLM, SGLang, TensorRT-LLM, CUDA, Triton, CUTLASS, PyTorch, TensorRT
Responsibilities
Implement and productionize cutting-edge inference techniques (quantization, speculative decoding, KV-cache reuse); Profile and optimize inference end-to-end from kernel launch to request scheduling; Build benchmarking frameworks for real-world performance; Bring up new model architectures on new hardware quickly; Contribute upstream to open-source inference engines. (via careerplan.io/jobs/7cb19a05-8e5b-44cf-b3e7-19949e2eff04-software-engineer-inference-performance-at-baseten)
Seniority
Senior, hands-on IC
