AML-机器学习系统研发工程师
Core
Develops and optimizes machine learning training and inference frameworks for ByteDance's AML platform, supporting internal business units (Douyin, Toutiao) and external clients via Volcano Engine.
Role type
Senior IC machine learning systems engineer (frameworks & infrastructure)
Builds
Volcano Engine ML platform, Fangzhou large model platform, and AI for Science research capabilities
Domain
Cloud computing, large-scale distributed systems, AI infrastructure
Deliverable
production ML models
Required skills
C/C++, Python, CUDA, Linux system programming, GDB, Nsight, distributed training, model inference optimization, resource scheduling, task orchestration, heterogeneous computing, AI compilers (MLIR/TVM/LLVM), NCCL/MPI/RDMA
Preferred skills
ACM/ICPC/Codeforces awards, deep learning frameworks (PyTorch/TensorFlow), large model architectures (Transformer/MoE/Sparse), quantization, RLHF, multi-modal models
Technologies
Linux, GDB, Nsight, NCCL, MPI, RDMA, DPDK, MLIR, TVM, Triton, LLVM, CUDA, OpenCL
Responsibilities
Develop and optimize ML training/inference frameworks; Solve high-concurrency, high-reliability, and scalability challenges; Cover multiple sub-domains including resource scheduling, task orchestration, and model management; Research and integrate cutting-edge hardware and heterogeneous computing technologies; Analyze and optimize cluster/service resource usage using ML methods.