CareerPlanGet AI match score →

Inference Engineer

San Francisco, CA🌐 Remote💼 Full-time🗓 2026-05-21 → 2026-07-29

Core

Building and optimising realtime TTS streaming infrastructure and runtime systems for low-latency conversational speech models to power enterprise voice experiences.

Role type

Senior IC machine-learning engineer (realtime inference systems)

Builds

Production-grade realtime speech inference systems serving hundreds of millions of conversations

Domain

Voice AI / Realtime speech infrastructure

Deliverable

production ML models

Required skills

Realtime inference optimisation, scheduler design, GPU utilisation, concurrency optimisation, dynamic batching, TensorRT, Triton, ONNX Runtime, CUDA, Rust, C++, Python, Kubernetes, vLLM, CUDA Graphs

Preferred skills

Experience with heterogeneous GPU environments (NVIDIA/AMD), speculative decoding, KV cache management, kernel-level profiling

Technologies

TensorRT, Triton, ONNX Runtime, vLLM, CUDA, Kubernetes, AWS, Rust, C++, Python

Responsibilities

Building and optimising realtime TTS streaming infrastructure, Improving scheduler and batching systems for production workloads, Reducing TTFA/TTFB while maintaining speech quality and stability, GPU profiling and identifying kernel-level bottlenecks, Optimising TensorRT, Triton, ONNX Runtime, and custom serving systems, Managing KV cache systems, speculative decoding, and streaming inference

Seniority

Senior, hands-on IC

Sourced via techire · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on techire.ai ↗