大模型训练调度专家-Seed
Core
Design and develop machine learning system resource scheduling to support model training, evaluation, and inference across NLP, CV, and Speech scenarios.
Role type
Senior IC machine learning systems engineer (resource scheduling)
Builds
Distributed ML training clusters and inference services for large language models and multimodal applications
Domain
Artificial Intelligence / Distributed Systems / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux environment development, Go/Python/Shell programming, Kubernetes architecture, Container technologies (Docker/Containerd/Kata/Podman), Distributed system design and maintenance, Resource optimization logic, RDMA network management, Storage resource scheduling, Multi-cloud orchestration, Task prioritization and preemption logic
Preferred skills
PyTorch/Megatron-LLM framework experience, Ray framework or reinforcement learning frameworks, Experience in data-driven ML systems, Large-scale AI task fault tolerance, High-performance computing (HPC), GPU hardware drivers, Publications in OSDI/SOSP/NSDI/ATC/EuroSys