大模型训练优化工程师 - Seed Model
Core
Design and develop large-scale machine learning system architectures to support the training of general-purpose large models (Seed Model), focusing on scalability, reliability, and ease of use.
Role type
Senior IC machine learning systems engineer (large model training)
Builds
Scalable, high-reliability distributed training systems for pretraining, reinforcement learning, and new hardware adaptation
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Large-scale system architecture design, distributed model training, high-performance computing, data management, resource scheduling, algorithm-system joint optimization
Preferred skills
LLM/NLP/CV/voice algorithms, Diffusion/RL algorithms, CUDA, torch.compile/Triton/TVM, RDMA, heterogeneous acceleration hardware, distributed systems, big data architecture
Technologies
CUDA, torch.compile, Triton, TVM, RDMA
Responsibilities
Design and develop large-scale ML system architectures; Research and implement cutting-edge training technologies; Collaborate with algorithm teams for joint optimization across pretraining, RL, and hardware adaptation; Manage sub-areas including distributed training, HPC, data management, and resource scheduling.