大模型在线服务SRE工程师(深圳/北京)
Core
Build and maintain SRE systems for ultra-large-scale general large language model services, ensuring stability and quality under high concurrency and complex heterogeneous resource environments.
Role type
Senior SRE Engineer (LLM Infra)
Builds
AI platform, GPU/CPU/network/storage infrastructure, and model inference services
Domain
AI Infrastructure / Large Language Models
Deliverable
production ML models
Required skills
Linux system tuning, Kubernetes, Python/Go programming, LLM training/inference workflows, capacity planning, fault diagnosis, resource utilization optimization, cloud-native technologies
Preferred skills
LLMOps/MLOps engineering, AI-native SRE concepts, intelligent on-call systems, infrastructure cost optimization, performance analysis
Technologies
Kubernetes, Docker, Python, Go, GPU clusters, distributed systems
Responsibilities
Design SRE systems for high-concurrency LLM services; Build monitoring, observability, and automated O&M platforms; Handle incident response and root cause analysis; Optimize resource utilization and inference costs; Participate in LLM platform deployment and tuning; Track and adopt emerging AI hardware and infrastructure technologies.
Seniority
Senior, hands-on IC