AML机器学习系统调度研发工程师-Data
Core
Design and develop machine learning system resource scheduling to support model training, evaluation, and inference for NLP, CV, and Speech scenarios.
Role type
Senior IC machine learning infrastructure engineer (scheduling)
Builds
Distributed ML training and inference clusters across multi-cloud and multi-region environments
Domain
AI Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
Linux, Go, Python, Shell, Kubernetes, Docker, Containerd, Kata, Podman, Distributed Systems, Resource Orchestration, RDMA, Storage Optimization
Preferred skills
TensorFlow, PyTorch, AI Infrastructure, HW/SW Co-Design, High Performance Computing, ML Hardware Architecture
Responsibilities
Design and develop resource scheduling systems for ML workloads; Optimize heterogeneous resource usage (GPU, CPU, multi-cloud); Manage distributed cluster compute, network, and storage resources; Distribute workloads across multi-datacenter and multi-region scenarios