Research Engineer, Audio and Speech
Core
Building multimodal and full-duplex AI voice agents that listen, reason, speak, and respond naturally in real-time for enterprise customer support.
Role type
Senior IC research engineer (audio/speech)
Builds
Real-time voice agents, streaming agent harnesses, and production ML models for conversational AI
Domain
Conversational AI, Speech Technology, Multimodal Systems
Deliverable
production ML models
Required skills
Speech/audio ML, multimodal ML, autoregressive/diffusion/flow-matching models, streaming agent systems, low-latency inference, production model serving, Python, deep learning frameworks (PyTorch), signal processing
Preferred skills
Speech-to-speech models, full-duplex models, telephony systems, multilingual speech, noisy-channel robustness, speaker adaptation, expressive speech generation
Technologies
PyTorch, autoregressive models, diffusion models, flow-matching models, codec-based models
Responsibilities
Design and build next-generation agent harnesses optimized for streaming speech and turn-taking; Research and train multimodal and full-duplex models; Improve speech recognition, voice activity detection, endpointing, and speech generation; Build evaluations and use production calls to ship measurable improvements; Optimize end-to-end inference for responsiveness, throughput, stability, and cost
Seniority
Senior, hands-on IC