CareerPlanSign in

IaaS研发实习生(J99325)

深圳市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Develop and maintain the core AI Infrastructure platform to support stable large model training and inference operations, managing GPU, RDMA, and high-speed network resources within a cloud-native environment.

Role type

IaaS Research Intern (AI Infrastructure)

Builds

Cloud-native AI infrastructure platform for large model training and inference

Domain

Cloud Computing / AI Infrastructure / Distributed Systems

Deliverable

production ML models

Required skills

Go, Python, C/C++, Linux, Kubernetes, Cloud Native, Distributed Systems, GPU management, RDMA, High-speed networking

Preferred skills

Large model training/inference workflows, Multi-card/multi-machine parallelism (TP/DP/PP/PD), HPC projects

Technologies

Kubernetes, Go, Python, C/C++, Linux, RDMA, GPU clusters

Responsibilities

Develop modules for GPU/RDMA resource management and scheduling; Build automated acceptance and performance testing systems for IaaS resources; Optimize Kubernetes container platform for stability and scalability; Implement platform engineering for large model training and inference; Resolve production issues in multi-node GPU clusters; Develop observability tools for capacity and cost management.

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.