CareerPlanSign in

Software Engineer - Voice AI (Inference Runtime)

San Francisco💼 Full-time🗓 2026-04-23 → 2026-09-26

Core

Primary owner of Baseten Voice AI in-house inference stack, building production-grade real-time systems for STT, TTS, and voice agent workloads to power mission-critical customer deployments.

Role type

Senior IC machine-learning infrastructure engineer (voice AI inference runtime)

Builds

Large-scale, real-time model serving systems for state-of-the-art open-source voice models

Domain

AI infrastructure / Voice AI / Real-time systems

Deliverable

production ML models

Required skills

System design, real-time large-scale system ownership, tail latency optimization, Python, cross-team collaboration, technical leadership, AI coding assistant proficiency

Preferred skills

Pipeline-level model runtime optimizations, developer platform building (SDKs, CLIs, APIs), containerization and orchestration, speech/audio ML models, model-serving runtimes, systems-level performance profiling

Technologies

vLLM, TensorRT, ONNX, Docker, Kubernetes, PyTorch Profiler

Responsibilities

Own and lead Voice AI product areas end-to-end from architecture to production operations; Design, build, and operate real-time, high-performance model serving systems; Drive cross-team collaboration to solve full-stack technical problems; Mentor teammates through code reviews and design docs

Seniority

Senior, hands-on IC with technical leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.