大模型训练优化工程师 - Seed Model
Core
Design and develop large-scale machine learning systems to optimize training for general-purpose AI models, ensuring scalability, reliability, and ease of use.
Role type
Senior IC machine learning systems engineer (large model training optimization)
Builds
Scalable, high-reliability distributed training infrastructure for pretraining, reinforcement learning, and new hardware adaptation
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed model training, High-performance computing, Data management, Resource scheduling, System architecture design, Algorithm-system joint optimization, CUDA programming, Compiler technologies (torch.compile/Triton/TVM), RDMA communication libraries, Heterogeneous hardware acceleration, Big data architecture
Preferred skills
LLM/NLP/CV/Speech algorithms, Diffusion algorithms, RL algorithms, System algorithm joint optimization
Responsibilities
Design and develop large-scale ML system architecture; Research and implement cutting-edge training technologies; Collaborate with algorithm teams for joint optimization across training scenarios; Manage sub-areas including distributed training, HPC, data management, and resource scheduling.