Audio AI Engineer
Core
Research and develop low-latency streaming algorithms for accent conversion, voice conversion, speech synthesis, and speech recognition to enhance real-time communication systems.
Role type
Senior IC Audio AI Engineer (Speech Processing)
Builds
Real-time audio models for Zoom's communication platform
Domain
Telecommunications / Speech AI
Deliverable
production ML models
Required skills
Deep learning frameworks (PyTorch, TensorFlow), Python/C++ programming, sequence modeling (Transformers, RNNs, diffusion models, conformers), low-latency model deployment, model compression (quantization, pruning, distillation), real-time audio system integration
Preferred skills
PhD or equivalent experience in streaming/ASR/TTS, 2+ years industry experience, publications in top-tier conferences (ICASSP, Interspeech, NeurIPS, ICLR)
Technologies
PyTorch, TensorFlow, Python, C/C++, Transformers, RNNs, diffusion models, conformers
Responsibilities
Research and design algorithms for accent/voice conversion and speech synthesis/recognition; prototype end-to-end audio models; integrate models into real-time video/audio systems; optimize model performance for quality, latency, and scalability; contribute to patents and knowledge sharing
Seniority
Senior, hands-on IC