CareerPlanSign in

魔方工作室-视频生成基础模型开发工程师

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Leading large-scale pretraining of video diffusion or autoregressive base models (DiT / Spatio-temporal Transformers) and optimizing them for real-time generation with control signals.

Role type

Senior IC machine-learning engineer (video generation)

Builds

Real-time video generation models with controllable outputs (first/last frame, trajectory, depth, mask, reference image)

Domain

AI / Computer Vision / Video Generation

Deliverable

production ML models

Required skills

Deep learning research, PyTorch, large-scale distributed training (FSDP/DeepSpeed/Megatron), diffusion models (DDPM/DDIM/Flow Matching), autoregressive generation, DiT/U-ViT architectures, video/image generation, model quantization (FP4/INT8), TensorRT, experimental design, paper reproduction

Preferred skills

None stated

Technologies

PyTorch, FSDP, DeepSpeed, Megatron-LM, TensorRT, FP4, INT8, DiT, U-ViT, DDPM, DDIM, Flow Matching, Rectified Flow

Responsibilities

Designing architecture and training objectives for video base models; Implementing control signal injection for identity and scene consistency; Optimizing models for real-time inference via causal forcing and distillation; Collaborating on inference optimization for latency and throughput; Building quantitative evaluation metrics for quality and consistency; Managing distributed training infrastructure on hundreds to thousands of GPUs

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 850,000+ jobs from 20+ sources.