CareerPlanSign in

机器学习平台调度工程师​​(北京/深圳/上海/杭州)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and optimize global resource scheduling for large-scale GPU clusters to ensure efficient and stable execution of offline and online AI training tasks.

Role type

Senior IC machine-learning platform scheduling engineer

Builds

High-availability scheduling frameworks supporting distributed training on large-scale GPU clusters

Domain

Cloud-native infrastructure for AI/ML training

Deliverable

production ML models

Required skills

Go, Python, C++, Kubernetes core components, RDMA, distributed storage, OpenMP, MPI, PyTorch, TensorFlow

Preferred skills

ARM heterogeneous computing, hybrid cloud, virtualization, CRD development, CSI plugins

Technologies

Kubernetes, Docker, RDMA, PyTorch, TensorFlow, OpenMP, MPI

Responsibilities

Lead global resource scheduling for 10,000+ GPU clusters; Optimize RDMA networks and distributed storage for large-scale training; Build high-availability scheduling frameworks using K8s; Develop K8s schedulers, CSI plugins, and CRDs for task orchestration and disaster recovery.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.