CareerPlanSign in

大模型在线服务SRE工程师(深圳/北京)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Build and maintain SRE systems for ultra-large-scale general large language model services, ensuring stability and quality under high concurrency and complex heterogeneous resource environments.

Role type

Senior SRE Engineer (LLM Infra)

Builds

AI platform, GPU/CPU/network/storage infrastructure, and model inference services

Domain

AI Infrastructure / Large Language Models

Deliverable

production ML models

Required skills

Linux system tuning, Kubernetes, Python/Go programming, LLM training/inference workflows, capacity planning, fault diagnosis, resource utilization optimization, cloud-native technologies

Preferred skills

LLMOps/MLOps engineering, AI-native SRE concepts, intelligent on-call systems, infrastructure cost optimization, performance analysis

Technologies

Kubernetes, Docker, Python, Go, GPU clusters, distributed systems

Responsibilities

Design SRE systems for high-concurrency LLM services; Build monitoring, observability, and automated O&M platforms; Handle incident response and root cause analysis; Optimize resource utilization and inference costs; Participate in LLM platform deployment and tuning; Track and adopt emerging AI hardware and infrastructure technologies.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.