CareerPlanSign in

Engineer, Inference & Model serving

San Francisco, CA💼 Full-time💰 $220,000–$220,000🗓 2026-04-30 → 2026-09-26

Core

Building high-performance, low-latency inference systems for LLMs, speech, and vision models in production.

Role type

Senior IC ML Model Serving Engineer

Builds

Real-time AI systems serving LLMs, speech, and vision models

Domain

Artificial Intelligence / Machine Learning Infrastructure

Deliverable

production ML models

Required skills

ML inference or model serving systems, Python, PyTorch, distributed systems, production infrastructure, latency and throughput optimisation, GPU profiling, CUDA, Kubernetes, Ray

Preferred skills

Systems or performance engineering mindset

Technologies

vLLM, TensorRT-LLM, Triton, SGLang, CUDA, Kubernetes, Ray

Responsibilities

Building high-performance serving systems for LLM, speech, and vision models, Scaling inference to production workloads with strict latency requirements, Optimising GPU utilisation and execution efficiency, Implementing techniques like continuous batching, KV cache optimisation, speculative decoding, and prefill/decode separation, Improving frameworks such as vLLM, TensorRT-LLM, Triton, and SGLang, Profiling and debugging performance across GPU, memory, and system layers

Seniority

Senior, hands-on IC

Sourced via techire · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.