魔方工作室-视频生成基础模型开发工程师
Core
Leading large-scale pretraining of video diffusion or autoregressive base models (DiT / Spatio-temporal Transformers) and optimizing them for real-time generation with control signals.
Role type
Senior IC machine-learning engineer (video generation)
Builds
Real-time video generation models with controllable outputs (first/last frame, trajectory, depth, mask, reference image)
Domain
AI / Computer Vision / Video Generation
Deliverable
production ML models
Required skills
Deep learning research, PyTorch, large-scale distributed training (FSDP/DeepSpeed/Megatron), diffusion models (DDPM/DDIM/Flow Matching), autoregressive generation, DiT/U-ViT architectures, video/image generation, model quantization (FP4/INT8), TensorRT, experimental design, paper reproduction
Preferred skills
None stated
Technologies
PyTorch, FSDP, DeepSpeed, Megatron-LM, TensorRT, FP4, INT8, DiT, U-ViT, DDPM, DDIM, Flow Matching, Rectified Flow
Responsibilities
Designing architecture and training objectives for video base models; Implementing control signal injection for identity and scene consistency; Optimizing models for real-time inference via causal forcing and distillation; Collaborating on inference optimization for latency and throughput; Building quantitative evaluation metrics for quality and consistency; Managing distributed training infrastructure on hundreds to thousands of GPUs
Seniority
Senior, hands-on IC