CareerPlanSign in

Research Engineer, Audio and Speech

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-09-04 → 2026-09-26

Core

Building multimodal and full-duplex AI voice agents that listen, reason, speak, and respond naturally in real-time for enterprise customer support.

Role type

Senior IC research engineer (audio/speech)

Builds

Real-time voice agents, streaming agent harnesses, and production ML models for conversational AI

Domain

Conversational AI, Speech Technology, Multimodal Systems

Deliverable

production ML models

Required skills

Speech/audio ML, multimodal ML, autoregressive/diffusion/flow-matching models, streaming agent systems, low-latency inference, production model serving, Python, deep learning frameworks (PyTorch), signal processing

Preferred skills

Speech-to-speech models, full-duplex models, telephony systems, multilingual speech, noisy-channel robustness, speaker adaptation, expressive speech generation

Technologies

PyTorch, autoregressive models, diffusion models, flow-matching models, codec-based models

Responsibilities

Design and build next-generation agent harnesses optimized for streaming speech and turn-taking; Research and train multimodal and full-duplex models; Improve speech recognition, voice activity detection, endpointing, and speech generation; Build evaluations and use production calls to ship measurable improvements; Optimize end-to-end inference for responsiveness, throughput, stability, and cost

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.