Staff Machine Learning Engineer, Siri Runtime Systems and Interaction
Core
Lead the development of audio and video generation capabilities for realistic, expressive synthetic speech and visual representations to enhance human-computer interaction.
Role type
Staff Machine Learning Engineer (Generative AI, Multimodal)
Builds
End-to-end systems for generating synthetic audio (speech, acoustics) and video for agent interactions
Domain
Generative AI, Speech Synthesis, Computer Vision
Deliverable
production ML models
Required skills
Generative audio/video architectures (diffusion, autoregressive, GANs, VAEs), Model evaluation and perceptual quality assessment, Technical architecture and tradeoffs, Mentoring engineers, Research-to-production leadership
Preferred skills
Speech synthesis (TTS), Voice conversion, Generative video/animation (facial animation, lip-sync), Multimodal modeling, Low-latency model deployment, Publications in top-tier venues
Technologies
Python, PyTorch, TensorFlow, Distributed training infrastructure