机器学习训练框架研发工程师 - Data AML
Core
Develop and maintain large-scale training systems for recommendation, advertising, and search models, supporting ultra-large sparse/dense/multi-modal/LLM training and exploring new paradigms like Scaling Law and LLM4Rec.
Role type
Senior IC machine-learning training framework engineer
Builds
Scalable training infrastructure for recommendation, ad, and search models serving internal business units and external enterprise clients via Volcano Engine.
Domain
Internet / Machine Learning Systems / Distributed Computing
Deliverable
production ML models
Required skills
C++, Python, CUDA, Linux, distributed systems, parallel training strategies, system architecture, performance optimization, stability governance
Preferred skills
PyTorch, DeepSpeed, Megatron-LM, Triton, GPU/NPU programming, NCCL/RDMA optimization, multi-modal representation learning, open-source contributions
Technologies
PyTorch, TensorFlow, JAX, Megatron-LM, DeepSpeed, FSDP, Ray, VeRL, CUDA, Triton, NCCL, RDMA, Kubernetes
Responsibilities
Design end-to-end training solutions for efficiency, stability, and cost; build distributed system modules (checkpoint, fault tolerance, observability); optimize GPU embedding, multi-level storage, and high-performance communication; collaborate across teams to upgrade training architecture.
Seniority
Senior, hands-on IC