CareerPlanSign in

北京-大模型训练Infra研发工程师(基座研发方向)(J101258)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Develop and optimize infrastructure for training large-scale foundation models (Baidu ERNIE) and the PaddlePaddle deep learning framework, focusing on distributed training architecture and high-performance computing.

Role type

Senior IC infrastructure engineer (large model training)

Builds

Distributed training systems, high-performance computing libraries, and engineering efficiency platforms for deep learning frameworks.

Domain

AI Infrastructure / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

C++, Python, CUDA programming, distributed training frameworks (PaddlePaddle, PyTorch, TensorFlow), Linux/Unix development, network programming, multi-threading

Preferred skills

Deep learning algorithm-engineering co-optimization, cluster fault tolerance, communication strategies (MPI, NCCL, RDMA, GPU Direct), hardware performance analysis (CodeXL, NVVP, GPA), cloud-native technologies (Kubernetes, Docker, Istio), parallel computing optimization

Technologies

CUDA, C++, Python, PaddlePaddle, PyTorch, TensorFlow, DeepSpeed, Megatron, NCCL, RDMA, Kubernetes, Docker, Istio, MPI, OpenBLAS, MKL, Eigen, Cublas, Cudnn, MIopen

Responsibilities

Optimize training efficiency for large models and PaddlePaddle; explore distributed training architectures; develop high-performance computing and communication libraries; maintain CI/CD pipelines and framework stability; build engineering efficiency platforms and automation tools.

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.