机器学习平台调度工程师(北京/深圳/上海/杭州)
Core
Design and optimize global resource scheduling for large-scale GPU clusters to ensure efficient and stable execution of offline and online AI training tasks.
Role type
Senior IC machine-learning platform scheduling engineer
Builds
High-availability scheduling frameworks supporting distributed training on large-scale GPU clusters
Domain
Cloud-native infrastructure for AI/ML training
Deliverable
production ML models
Required skills
Go, Python, C++, Kubernetes core components, RDMA, distributed storage, OpenMP, MPI, PyTorch, TensorFlow
Preferred skills
ARM heterogeneous computing, hybrid cloud, virtualization, CRD development, CSI plugins
Technologies
Kubernetes, Docker, RDMA, PyTorch, TensorFlow, OpenMP, MPI
Responsibilities
Lead global resource scheduling for 10,000+ GPU clusters; Optimize RDMA networks and distributed storage for large-scale training; Build high-availability scheduling frameworks using K8s; Develop K8s schedulers, CSI plugins, and CRDs for task orchestration and disaster recovery.
Seniority
Senior, hands-on IC