CareerPlanSign in

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

San Francisco, CA💼 Full-time💰 $170,000–$170,000🗓 2026-05-08 → 2026-09-26

Core

Building high-throughput, ultra-low-latency inference engines for large language models and foundational speech models to power real-time conversational AI.

Role type

Senior IC machine learning engineer (inference & serving)

Builds

Real-time speech LLM inference systems for hardware-software AI companions

Domain

AI/ML, Speech Technology, Distributed Systems

Deliverable

production ML models

Required skills

GPU architecture optimization, continuous batching, KV cache management, real-time audio streaming, model compression & quantization, distributed inference pipelines

Preferred skills

vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, speculative decoding, chunked prefill, Kubernetes autoscaling

Technologies

NVIDIA Ampere/Hopper GPUs, WebSockets, WebRTC, Kubernetes, FP8, INT8, AWQ, GPTQ

Responsibilities

Deploy multi-GPU and multi-node inference pipelines, manage autoscaling infrastructure, optimize hardware bottlenecks, implement advanced generation algorithms, handle continuous audio streams

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.