分布式计算研发工程师(J92969)
Core
Design and develop AI R&D platforms and heterogeneous computing foundations for large-scale AI computing clusters (100k+ GPU cards) to maximize compute utilization and efficiency.
Role type
Senior IC distributed systems engineer (AI infrastructure)
Builds
Serverless heterogeneous compute platforms, cloud-native AI components, and solutions for development, training, and inference scenarios.
Domain
AI Infrastructure / Distributed Systems / Cloud Native
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals (scheduler, Device Plugin, CRI, CNI), Go/Python/C++, large-scale distributed system design, GPU cluster management, high-performance networking, AI storage systems
Preferred skills
Kubeflow, Volcano, Ray, CUDA, NVLink, InfiniBand, complex system debugging
Responsibilities
Design heterogeneous multi-core compute platforms, optimize GPU virtualization and intelligent over-provisioning, integrate AI storage and high-performance networks, evolve system architecture for stability and scalability