CareerPlanSign in

机器学习平台SRE工程师(深圳/北京)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Define SLA/SLO, build monitoring/alerting, and ensure stability of the machine learning platform's training and inference chains.

Role type

Senior SRE for Machine Learning Platform

Builds

AI training pipelines, resource scheduling systems, and storage infrastructure for ML workloads

Domain

Artificial Intelligence / Distributed Systems / Cloud Infrastructure

Deliverable

infrastructure

Required skills

Kubernetes, Python, Go, Shell, distributed system operations, high-concurrency system design, database performance tuning (MySQL/PostgreSQL), GPU cluster management, NCCL/RDMA/InfiniBand networking, LLM training/inference optimization, PyTorch, DeepSpeed, vLLM, SGLang, Agent/Harness architecture

Preferred skills

Experience with AI-driven automation for operational tasks, MCP (Model Context Protocol) integration

Responsibilities

Define SLA/SLO and manage fault response for the ML platform; Architect disaster recovery and performance tuning for the platform; Automate operational workflows using AI capabilities for resource and change management.

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.