Software Engineer - Voice AI (Inference Runtime)
Core
Primary owner of Baseten Voice AI in-house inference stack, building production-grade real-time systems for STT, TTS, and voice agent workloads to power mission-critical customer deployments.
Role type
Senior IC machine-learning infrastructure engineer (voice AI inference runtime)
Builds
Large-scale, real-time model serving systems for state-of-the-art open-source voice models
Domain
AI infrastructure / Voice AI / Real-time systems
Deliverable
production ML models
Required skills
System design, real-time large-scale system ownership, tail latency optimization, Python, cross-team collaboration, technical leadership, AI coding assistant proficiency
Preferred skills
Pipeline-level model runtime optimizations, developer platform building (SDKs, CLIs, APIs), containerization and orchestration, speech/audio ML models, model-serving runtimes, systems-level performance profiling
Technologies
vLLM, TensorRT, ONNX, Docker, Kubernetes, PyTorch Profiler
Responsibilities
Own and lead Voice AI product areas end-to-end from architecture to production operations; Design, build, and operate real-time, high-performance model serving systems; Drive cross-team collaboration to solve full-stack technical problems; Mentor teammates through code reviews and design docs
Seniority
Senior, hands-on IC with technical leadership