机器学习平台SRE工程师(深圳/北京)
Core
Define SLA/SLO, build monitoring/alerting, and ensure stability of the machine learning platform's training and inference chains.
Role type
Senior SRE for Machine Learning Platform
Builds
AI training pipelines, resource scheduling systems, and storage infrastructure for ML workloads
Domain
Artificial Intelligence / Distributed Systems / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes, Python, Go, Shell, distributed system operations, high-concurrency system design, database performance tuning (MySQL/PostgreSQL), GPU cluster management, NCCL/RDMA/InfiniBand networking, LLM training/inference optimization, PyTorch, DeepSpeed, vLLM, SGLang, Agent/Harness architecture
Preferred skills
Experience with AI-driven automation for operational tasks, MCP (Model Context Protocol) integration
Responsibilities
Define SLA/SLO and manage fault response for the ML platform; Architect disaster recovery and performance tuning for the platform; Automate operational workflows using AI capabilities for resource and change management.