CareerPlanSign in

Software Engineer- Inference Performance

San Francisco💼 Full-time🗓 2026-10-05 → 2026-10-07

Core

Build and optimize the inference engine and runtime to make demanding AI workloads run faster and more efficiently for customers.

Role type

Senior IC inference performance engineer (LLM)

Builds

High-performance inference stack, runtime internals, and scheduling systems for LLMs

Domain

AI/ML infrastructure, Large Language Models (LLM), GPU computing

Deliverable

production ML models

Required skills

C++, Python, LLM optimization techniques, PyTorch, TensorRT, GPU architecture, speculative decoding, quantization, KV-cache management, cross-layer profiling

Preferred skills

CUDA/Triton kernel development, upstream open-source contributions (vLLM, SGLang), large-scale distributed serving, FP8/FP4 quantization

Technologies

vLLM, SGLang, TensorRT-LLM, CUDA, Triton, CUTLASS, PyTorch, TensorRT

Responsibilities

Implement and productionize cutting-edge inference techniques (quantization, speculative decoding, KV-cache reuse); Profile and optimize inference end-to-end from kernel launch to request scheduling; Build benchmarking frameworks for real-world performance; Bring up new model architectures on new hardware quickly; Contribute upstream to open-source inference engines. (via careerplan.io/jobs/7cb19a05-8e5b-44cf-b3e7-19949e2eff04-software-engineer-inference-performance-at-baseten)

Seniority

Senior, hands-on IC