GPU算力平台研发实习生(J106052)
Core
Design and develop infrastructure and products for large-scale AI computing clusters, focusing on heterogeneous multi-chip cluster construction and GPU resource optimization.
Role type
GPU computing platform R&D intern
Builds
Heterogeneous computing platforms and solutions for development, training, and inference scenarios
Domain
AI infrastructure, cloud-native systems, distributed computing
Deliverable
production ML models
Required skills
Kubernetes development, container runtime, container networking, GPU chip architecture, distributed system architecture
Preferred skills
Kubeflow, Volcano, PyTorch
Technologies
Kubernetes, Kubeflow, Volcano, PyTorch
Responsibilities
Design and develop cloud-native AI components including training/inference orchestration, GPU scheduling, and high-performance networking; Optimize distributed system architecture for stability, performance, and scalability.
Seniority
Intern