AI Engineer, Voice Designer
Core
Own the back-end implementation and linguistic optimization of the Text-to-Speech (TTS) layer for next-generation AI voice agents, ensuring they sound human and context-aware.
Role type
Senior IC machine-learning engineer (voice/TTS)
Builds
Production-grade TTS systems, voice personas, and conversational fillers for agentic AI agents
Domain
Business communications, AI speech synthesis, agentic AI
Deliverable
production ML models
Required skills
Python, deep learning frameworks (PyTorch), Speech Synthesis (TTS), phonetics, sociolinguistics, SSML orchestration, LLM prompt engineering, backend API development, speech quality metrics (MOS, latency)
Preferred skills
Experience with NVIDIA NeMo, ESPnet, Coqui, ElevenLabs, Rime, Cartesia, GCP
Technologies
PyTorch, NVIDIA NeMo, ESPnet, Coqui, ElevenLabs, Rime, Cartesia, GCP
Responsibilities
Integrate and optimize multiple TTS vendor APIs; apply phonetics and sociolinguistics for natural TTS input; craft context-specific utterances for turn handling; design and manage LLM/TTS prompts for agent personalities; architect UI logic for voice attribute customization; partner with ASR and Audio AI engineers for end-to-end voice quality
Seniority
Senior, hands-on IC