云原生算力平台SRE工程师(深圳/北京/上海)
Core
Operate, troubleshoot, and optimize GPU/CPU heterogeneous computing infrastructure to ensure stable, efficient, and continuous compute output.
Role type
Senior Site Reliability Engineer (Compute Platform)
Builds
High-availability GPU/CPU compute clusters and automated operations tooling
Domain
Cloud-native computing infrastructure and heterogeneous hardware
Deliverable
production ML models | infrastructure
Required skills
GPU hardware and driver tuning, Kubernetes cluster management, Linux administration, Golang/Python/Java, heterogeneous computing (Mellanox/NCCL/Cuda)
Preferred skills
Experience with Docker, cloud-native disaster recovery design, automation and intelligent operations methods
Technologies
Kubernetes, Docker, Golang, Python, Java, Linux, CUDA, NCCL, Mellanox
Responsibilities
Daily operations and troubleshooting of GPU/CPU heterogeneous devices; Managing and governing K8s clusters including disaster recovery and security drills; Automating operational workflows for resource and change management.