高级SRE工程师(J94322)
Core
Ensure stability, availability, and performance for a large language model (LLM) business; manage GPU/CPU infrastructure capacity and cost; build observability and automated operations platforms.
Role type
Senior Site Reliability Engineer (LLM Infrastructure)
Builds
LLM service infrastructure, automated operations toolchains, observability systems
Domain
Cloud Infrastructure, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
Linux system administration, GPU architecture and CUDA, Kubernetes, Golang or Python, distributed system resource management, fault diagnosis
Preferred skills
None stated
Technologies
Kubernetes, CUDA, Golang, Python
Responsibilities
Monitor and resolve system faults, optimize resource usage efficiency, design and develop automation tools
Seniority
Senior, hands-on IC
Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.