CareerPlanGet AI match score →

Research MLE (Training Optimization)大模型训练优化工程师

Beijing, Beijing, cn💼 Full-time🗓 2026-07-13 → 2026-07-31

Core

Design, implement, and optimize large-scale distributed training systems for multimodal and foundation models to improve GPU utilization, communication overhead, and memory efficiency.

Role type

Senior IC machine learning engineer (training optimization)

Builds

Distributed training infrastructure for large-scale multimodal and foundation models

Domain

Generative AI, Large Language Models, Multimodal AI

Deliverable

production ML models

Required skills

Distributed training systems, GPU optimization, CUDA/Triton kernel development, PyTorch, Megatron-LM, NeMo, FSDP/ZeRO, gradient checkpointing, low-precision data types, Python, C++ or Rust

Preferred skills

Experience with diffusion models, system programming languages

Technologies

Megatron-LM, NVIDIA NeMo, FSDP, Triton, CUDA, PyTorch, DeepSpeed

Responsibilities

Design and optimize large-scale ML training systems, improve performance across compute/memory/communication layers, partner with research teams, debug and profile training workflows, write custom GPU kernels

Seniority

Senior, hands-on IC

Sourced via smartrecruiters · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on SmartRecruiters ↗