大模型训练调度专家-Seed
Core
Design and develop machine learning system resource scheduling to support model training, evaluation, and inference across NLP, CV, and Speech scenarios.
Role type
Senior IC machine learning system engineer (resource scheduling)
Builds
Distributed ML training clusters and inference services for MLLM, GenMedia, and AI for Science applications
Domain
AI Infrastructure / Distributed Systems / Cloud Computing
Deliverable
infrastructure
Required skills
Linux environment development, Go/Python/Shell programming, Kubernetes architecture, Container technologies (Docker/Containerd/Kata/Podman), Distributed system design, Resource optimization (GPU/CPU/Heterogeneous), RDMA networking, Multi-cloud orchestration
Preferred skills
ML frameworks (TensorFlow/PyTorch), AI Infrastructure, HW/SW Co-Design, High Performance Computing, ML Hardware Architecture
Technologies
Kubernetes, Docker, Containerd, Kata, Podman, RDMA, TensorFlow, PyTorch, Go, Python, Shell, Linux
Responsibilities
Design and develop ML system resource scheduling for training, evaluation, and inference; Optimize heterogeneous resource orchestration (GPU, CPU, mixed, multi-cloud); Manage compute, RDMA network, and storage resource scheduling; Handle on-premise and cloud task/service scheduling across multiple data centers and regions.