CareerPlanSign in

机器学习训练框架研发工程师 - Data AML

北京💼 Full-time🗓 2026-09-28

Core

Develop and maintain large-scale training systems for recommendation, advertising, and search models, supporting ultra-large sparse/dense/multi-modal/LLM training and exploring new paradigms like Scaling Law and LLM4Rec.

Role type

Senior IC machine-learning training framework engineer

Builds

Scalable training infrastructure for recommendation, ad, and search models serving internal business units and external enterprise clients via Volcano Engine.

Domain

Internet / Machine Learning Systems / Distributed Computing

Deliverable

production ML models

Required skills

C++, Python, CUDA, Linux, distributed systems, parallel training strategies, system architecture, performance optimization, stability governance

Preferred skills

PyTorch, DeepSpeed, Megatron-LM, Triton, GPU/NPU programming, NCCL/RDMA optimization, multi-modal representation learning, open-source contributions

Technologies

PyTorch, TensorFlow, JAX, Megatron-LM, DeepSpeed, FSDP, Ray, VeRL, CUDA, Triton, NCCL, RDMA, Kubernetes

Responsibilities

Design end-to-end training solutions for efficiency, stability, and cost; build distributed system modules (checkpoint, fault tolerance, observability); optimize GPU embedding, multi-level storage, and high-performance communication; collaborate across teams to upgrade training architecture.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 849,000+ jobs from 20+ sources.