大模型后训练优化工程师 - Seed Model
Core
Design and develop ultra-large-scale machine learning system architectures to support pretraining, reinforcement learning, and new hardware adaptation for foundational models.
Role type
Senior IC machine learning systems engineer (LLM infrastructure)
Builds
Ultra-large-scale training systems and foundational model frameworks for products like Doubao and Jiemeng
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
ultra-large-scale system architecture design, distributed systems, algorithm-system joint optimization, low-precision/compression/matrix decomposition, heterogeneous acceleration hardware, RDMA/communication libraries, storage and I/O, CUDA, system scalability, high reliability
Preferred skills
LLM/NLP/CV/voice algorithms, Diffusion/RL algorithms, complete algorithm R&D and training flow, comprehensive system design and solution planning
Responsibilities
Design and develop ultra-large-scale ML system architectures, conduct research and implement forward-looking training solutions, collaborate with algorithm teams for joint optimization, update and refactor ML base frameworks and iteration scaffolds
