CareerPlanSign in

大模型训练调度专家-Seed

上海💼 Full-time🗓 2026-09-28

Core

Design and develop machine learning system resource scheduling to support model training, evaluation, and inference across NLP, CV, and Speech scenarios.

Role type

Senior IC machine learning systems engineer (resource scheduling)

Builds

Distributed ML training clusters and inference services for large language models and multimodal applications

Domain

Artificial Intelligence / Distributed Systems / Cloud Infrastructure

Deliverable

infrastructure

Required skills

Linux environment development, Go/Python/Shell programming, Kubernetes architecture, Container technologies (Docker/Containerd/Kata/Podman), Distributed system design and maintenance, Resource optimization logic, RDMA network management, Storage resource scheduling, Multi-cloud orchestration, Task prioritization and preemption logic

Preferred skills

PyTorch/Megatron-LLM framework experience, Ray framework or reinforcement learning frameworks, Experience in data-driven ML systems, Large-scale AI task fault tolerance, High-performance computing (HPC), GPU hardware drivers, Publications in OSDI/SOSP/NSDI/ATC/EuroSys

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.