机器学习系统调度研发工程师-Data AML
Core
Design and develop machine learning system resource scheduling to support model training, evaluation, and inference across NLP, CV, and Speech scenarios.
Role type
Senior IC machine learning system scheduling engineer
Builds
Optimized scheduling for heterogeneous resources (GPU, CPU, RDMA, storage) across multi-cloud, multi-region, and multi-datacenter environments
Domain
AI Infrastructure / Distributed Systems / Cloud Computing
Deliverable
infrastructure
Required skills
Linux environment development, Go/Python/Shell programming, Kubernetes architecture, Container technologies (Docker/Containerd/Kata/Podman), Distributed system design and development
Preferred skills
Machine learning frameworks (TensorFlow/Pytorch), AI Infrastructure, HW/SW Co-Design, High Performance Computing, ML Hardware Architecture
Technologies
Kubernetes, Docker, Containerd, Kata, Podman, RDMA, TensorFlow, PyTorch