大模型后训练优化工程师 - Seed Model
Core
Design and develop ultra-large-scale machine learning system architectures to support pretraining, reinforcement learning, and new hardware adaptation for foundational models.
Role type
Senior IC machine learning systems engineer (LLM infrastructure)
Builds
Ultra-large-scale training systems and foundational model frameworks
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
infrastructure
Required skills
Ultra-large-scale system architecture design, High-performance computing (HPC), CUDA programming, Distributed systems, Low-precision/compression/matrix factorization, Storage and I/O optimization, Heterogeneous hardware acceleration, System-algorithm joint optimization, Complex technical problem solving
Preferred skills
LLM/NLP/CV/voice algorithms, Diffusion/RL algorithms, Full algorithm R&D and training lifecycle experience, Engineering management and process optimization
Technologies
CUDA, RDMA, Diffusion, RL, LLM, NLP, CV
Responsibilities
Design and develop ultra-large-scale ML system architectures; Research and implement forward-looking training solutions; Collaborate with algorithm teams for joint optimization; Update and refactor ML base frameworks and iteration scaffolds.
