资源规划管理专家-Data AML
Core
Global compute resource supply-demand balancing and long-term evolution planning for AIGC/GPU-driven workloads.
Role type
Senior IC resource planning and capacity management expert
Builds
Global resource equilibrium strategies, data center/network/server roadmaps, and a self-service resource management platform
Domain
Cloud infrastructure, AIGC, GPU computing, and FinOps
Deliverable
production ML models
Required skills
Global resource planning, capacity management, SRE, FinOps, heterogeneous compute resource pool management, GPU technology stack (A100/H100/B200), container orchestration (Kubernetes), multi-tenancy strategies, mathematical modeling, strategic communication
Preferred skills
Experience with large-scale GPU clusters, cross-team leadership
Technologies
Kubernetes, NVLink, NVSwitch, RDMA
Responsibilities
Build global supply-demand prediction models, optimize physical infrastructure layout, design the resource self-service platform architecture, manage quota allocation workflows, embed capacity views and health metrics, and align data center geography with business trends.
Seniority
Senior, hands-on IC