大模型推理调度编排专家 - Seed Model
Core
Design and implement scheduling and orchestration systems for massive-scale LLM inference across heterogeneous resources to maximize cluster efficiency and stability.
Role type
Senior IC LLM inference scheduling and orchestration engineer
Builds
Scalable, multi-tenant LLM inference services supporting 50+ applications (e.g., Doubao, Jimeng, TRAE) via Volcano Engine
Domain
AI Infrastructure / Distributed Systems / Cloud Computing
Deliverable
production ML models
Required skills
C++/Go/Python/Shell, Kubernetes, Docker/Containerd/Kata/Podman, Distributed Systems, RDMA networking, GPU system architecture, Multi-cloud orchestration, Fault diagnosis and recovery
Preferred skills
vLLM/SGLang/PyTorch, LLM resource scheduling experience, Top-tier systems conference publications (OSDI/NSDI/SOSP/FAST/Eurosys)
Responsibilities
Manage heterogeneous resource scheduling, compute pooling, elastic scaling, and Quota management; Implement multi-role, multi-stage PD/EP scheduling and KVCache-centric dynamic scaling; Optimize compute, RDMA, and storage resource orchestration for distributed clusters; Ensure service stability through cross-system diagnostics and recovery in multi-cloud environments; Distribute workloads across multi-datacenter and multi-region scenarios
Seniority
Senior, hands-on IC