IaaS研发实习生(J99325)
Core
Develop and maintain the core AI Infrastructure platform to support stable large model training and inference operations, managing GPU, RDMA, and high-speed network resources within a cloud-native environment.
Role type
IaaS Research Intern (AI Infrastructure)
Builds
Cloud-native AI infrastructure platform for large model training and inference
Domain
Cloud Computing / AI Infrastructure / Distributed Systems
Deliverable
production ML models
Required skills
Go, Python, C/C++, Linux, Kubernetes, Cloud Native, Distributed Systems, GPU management, RDMA, High-speed networking
Preferred skills
Large model training/inference workflows, Multi-card/multi-machine parallelism (TP/DP/PP/PD), HPC projects
Technologies
Kubernetes, Go, Python, C/C++, Linux, RDMA, GPU clusters
Responsibilities
Develop modules for GPU/RDMA resource management and scheduling; Build automated acceptance and performance testing systems for IaaS resources; Optimize Kubernetes container platform for stability and scalability; Implement platform engineering for large model training and inference; Resolve production issues in multi-node GPU clusters; Develop observability tools for capacity and cost management.