高级 IaaS 研发工程师(J97206)
Core
Design and build the core AI Infrastructure platform to support stable and efficient large model training and inference operations, managing GPU, RDMA, and high-speed network resources within a cloud-native system.
Role type
Senior IC IaaS Platform Engineer (Cloud Native)
Builds
Cloud-native AI infrastructure platform for large model training and inference
Domain
Cloud Computing / AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Go, Kubernetes, Linux internals, Cloud-native system design, GPU resource management, RDMA, Performance analysis, Multi-tenant isolation, Distributed system debugging
Preferred skills
GPU cluster engineering, High-performance computing (HPC), Large model training/inference workflows, Parallel computing patterns (TP/DP/PP/PD)
Technologies
Kubernetes, Go, Linux, RDMA, GPU
Responsibilities
Design and evolve the Kubernetes container platform architecture for high availability and scalability; Automate GPU/RDMA resource admission, capability identification, and benchmarking; Support engineering implementation of large model training and inference on the platform; Manage resource scheduling and stability for multi-node GPU clusters; Build observability systems for capacity management and cost governance.
Seniority
Senior, hands-on IC