北京-大模型训练Infra研发工程师(基座研发方向)(J101258)
Core
Develop and optimize infrastructure for training large-scale foundation models (Baidu ERNIE) and the PaddlePaddle deep learning framework, focusing on distributed training architecture and high-performance computing.
Role type
Senior IC infrastructure engineer (large model training)
Builds
Distributed training systems, high-performance computing libraries, and engineering efficiency platforms for deep learning frameworks.
Domain
AI Infrastructure / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
C++, Python, CUDA programming, distributed training frameworks (PaddlePaddle, PyTorch, TensorFlow), Linux/Unix development, network programming, multi-threading
Preferred skills
Deep learning algorithm-engineering co-optimization, cluster fault tolerance, communication strategies (MPI, NCCL, RDMA, GPU Direct), hardware performance analysis (CodeXL, NVVP, GPA), cloud-native technologies (Kubernetes, Docker, Istio), parallel computing optimization
Technologies
CUDA, C++, Python, PaddlePaddle, PyTorch, TensorFlow, DeepSpeed, Megatron, NCCL, RDMA, Kubernetes, Docker, Istio, MPI, OpenBLAS, MKL, Eigen, Cublas, Cudnn, MIopen
Responsibilities
Optimize training efficiency for large models and PaddlePaddle; explore distributed training architectures; develop high-performance computing and communication libraries; maintain CI/CD pipelines and framework stability; build engineering efficiency platforms and automation tools.
Seniority
Senior, hands-on IC