资源调度与计算效率研发工程师/架构师 - AI算力基础设施
Core
Design and build a unified resource scheduling system for ByteDance's AI-native cloud infrastructure, managing GPU/CPU/memory/network resources for LLM training, inference, and online services.
Role type
Senior IC resource scheduling engineer/architect (AI infrastructure)
Builds
Scalable resource scheduling platform, AI elastic resource pools, and gang scheduling capabilities for LLM workloads
Domain
Cloud infrastructure, AI computing, distributed systems
Deliverable
production ML models | product features
Required skills
Distributed systems, system design, Go/C++/Java/Rust, Kubernetes/Yarn/Mesos/Slurm, cluster management, resource optimization, data analysis
Preferred skills
GPU scheduling, hardware topology (NUMA/NVLink), AI training workloads, operations research, reinforcement learning, open-source contributions
Technologies
Kubernetes, Slurm, CUDA, NCCL, RDMA, PyTorch, TensorFlow, Linux kernel
Responsibilities
Design unified resource expression and quota mechanisms for diverse workloads; Solve performance, availability, and fragmentation issues in large-scale schedulers; Implement GPU topology affinity and heterogeneous card adaptation; Build resource profiling and intelligent strategy systems; Collaborate with SRE and hardware teams on complex projects
Seniority
Senior, hands-on IC