云交付研发工程师(J90647)
Core
Design and develop technical solutions for scalable cloud operations, focusing on server lifecycle management, efficient delivery, and fault prediction for large-scale distributed systems.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure & AI Ops)
Builds
Intelligent cloud services, GPU hardware optimization mechanisms, and LLM-powered intelligent operations products.
Domain
Cloud Computing, Distributed Systems, GPU Hardware, Large Language Models
Deliverable
production ML models
Required skills
Linux system administration, Python, Go, Shell, GPU architecture knowledge, containerization (k8s), distributed storage, virtualization networks, KVM, OpenStack
Preferred skills
Deep learning and model training experience, hardware component principles, LLM fine-tuning and inference service construction
Responsibilities
Design intelligent cloud operations solutions, optimize cloud delivery quality assurance strategies, analyze GPU hardware faults, develop LLM-based operations tools, research industry technologies.
Seniority
Mid-Senior, hands-on IC
