大模型训练系统与优化工程师(VLM/Agent RL方向)-Data
Core
Designing and optimizing large-scale distributed training systems and frameworks for post-training, reinforcement learning, and multi-modal models.
Role type
Senior IC machine-learning systems engineer (LLM training infrastructure)
Builds
Scalable training frameworks for 100B+ parameter models, Agent RL harnesses, and multi-modal model architectures.
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Python, C++, PyTorch, DeepSpeed, Megatron, FSDP, distributed training optimization, operator fusion, memory optimization, RL training frameworks (PPO/GRPO), convergence debugging
Preferred skills
Agent RL framework development, Agentic Harness construction, MoE/Linear Attention architecture support, multi-modal training validation
Responsibilities
Abstracting and refactoring post-training frameworks for multi-modal compatibility, implementing distributed training strategies (DP/TP/PP/EP) for 100B~1T parameter models, building standardized evaluation benchmarks and stable harnesses for Agent RL, supporting novel model structures and multi-modal training convergence
Seniority
Senior, hands-on IC