Research Staff, Data Science
Core
Building an industrial 'data factory' to power next-generation Voice AI systems, focusing on data acquisition, preparation, synthesis, and advanced characterization of conversational audio.
Role type
Senior IC research data scientist (audio/voice AI)
Builds
Production-grade data pipelines, datasets, and benchmarking methodologies for speech and language AI foundation models
Domain
Voice AI, audio signal processing, deep learning
Deliverable
production ML models
Required skills
Data pipeline architecture, statistical methods, deep learning, Python, PyTorch, audio signal processing, data synthesis, automated labeling systems
Preferred skills
Physics, mechanical engineering, language processing background, model building experience, speech and audio domain expertise
Technologies
Python, PyTorch
Responsibilities
Drive high-performance data acquisition, preparation, and synthesis pipelines; Develop advanced characterizations of complex conversational audio; Collaborate with DataOps to create automated labeling systems; Build benchmarking methodologies and curated datasets; Document and present data experiment results
Seniority
Senior, hands-on IC