公有云机器学习系统工程师-调度方向
Core
Design and develop machine learning system resource scheduling to support large model platforms and ML product business.
Role type
Senior IC machine learning systems engineer (scheduling)
Builds
Optimized resource orchestration for multi-cluster, multi-tenant environments supporting offline training and online inference.
Domain
Cloud-native ML infrastructure, distributed systems, heterogeneous computing
Deliverable
production ML models
Required skills
Go/Java/Python, Linux, data structures and algorithms, ML frameworks (TensorFlow/PyTorch), Kubernetes, Docker/Containerd/Kata, distributed systems principles, system abstraction
Preferred skills
Large-scale cluster scheduling (K8S/Volcano/Yarn/Mesos), scheduling algorithms (Quota, preemption, elasticity, fragmentation, co-scheduling, QoS), GPU scheduling, CUDA, RDMA, AI Infrastructure, HW/SW Co-Design, High Performance Computing, Distributed Storage
Responsibilities
Design and develop resource scheduling for heterogeneous compute/storage/network resources; optimize resource utilization across multi-tenant isolation environments; support various load scenarios including offline training and online inference.
Seniority
Senior, hands-on IC