CareerPlanSign in

高级SRE工程师(J94322)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Ensure stability, availability, and performance for a large language model (LLM) business; manage GPU/CPU infrastructure capacity and cost; build observability and automated operations platforms.

Role type

Senior Site Reliability Engineer (LLM Infrastructure)

Builds

LLM service infrastructure, automated operations toolchains, observability systems

Domain

Cloud Infrastructure, Large Language Models, Distributed Systems

Deliverable

production ML models

Required skills

Linux system administration, GPU architecture and CUDA, Kubernetes, Golang or Python, distributed system resource management, fault diagnosis

Preferred skills

None stated

Technologies

Kubernetes, CUDA, Golang, Python

Responsibilities

Monitor and resolve system faults, optimize resource usage efficiency, design and develop automation tools

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.