Staff Research Engineer - Multimodal Generative Modelling
Core
Building human-interactive multimodal models that combine text, audio, and video for real-time conversational systems in enterprise skill development and visual communication.
Role type
Staff Research Engineer (Multimodal Generative Modelling)
Builds
Real-time interactive voice-video synthesis systems for enterprise customers
Domain
AI, Generative Models, Multimodal Systems, Audio/Video/Speech
Deliverable
production ML models
Required skills
Generative modelling, Large language models (LLMs), Transformer architectures, PyTorch, Time-series modeling, Tokenization, Deep learning model training, Software engineering
Preferred skills
Streaming architectures, Diffusion models, Neural codecs, Flow-matching models, Autoregressive decoders, Original research contributions
Technologies
PyTorch, Diffusion, Neural codecs, Flow-matching, Autoregressive models, LLMs
Responsibilities
Shape roadmap for new model capabilities, Propose novel multi-modal system architectures, Develop and evaluate streaming conversational systems, Design solutions for emotional expressiveness, Implement pretraining through post-training, Integrate and test novel architectures, Define new evaluation metrics, Curate datasets, Lead post-training initiatives, Ship models to production
Seniority
Staff, hands-on IC with strategic scope