Senior Research Engineer - Video Foundation Models (Pre - Training)
Core
Building production-grade foundation models for human-centric video generation, specifically latent video diffusion models for synthetic humans.
Role type
Senior Research Engineer (Applied Research)
Builds
Production-scale video base models powering realistic, controllable, and emotionally expressive synthetic humans
Domain
Generative AI, Video Foundation Models, Distributed Systems
Deliverable
production ML models
Required skills
Deep learning model training at scale, Python, PyTorch, Diffusion models (image/video), Large-scale multi-GPU/multi-node training, Distributed training (DDP, FSDP, DeepSpeed), Experimental design and analysis
Preferred skills
Video diffusion models, Avatar/human-centric generation, World/interactive models, GANs/VAEs, Inference system optimization
Technologies
Python, PyTorch, CUDA, DeepSpeed, Sequence parallelism, AWS, SLURM, Docker, GitHub, CI/CD
Responsibilities
Develop and scale latent video diffusion models, Design conditioning mechanisms for pose/emotion/script control, Advance distributed training strategies, Improve training stability at multi-node scale, Design rigorous evaluation frameworks, Optimize inference for low latency and cost efficiency, Run controlled ablations and experiments
Seniority
Senior, hands-on IC