CareerPlanSign in

机器学习平台SRE工程师-AML

上海💼 Full-time🗓 2026-09-28

Core

Ensure stability of machine learning systems supporting model development, training, and deployment; manage GPU/NPU/CPU and storage resources; handle multi-region disaster recovery and cluster governance.

Role type

Senior SRE for Machine Learning Platforms

Builds

Automated tools and platforms for resource management and operational efficiency

Domain

Cloud Infrastructure / Machine Learning Operations

Deliverable

infrastructure

Required skills

Linux administration, Go, Python, Shell scripting, Kubernetes, distributed system resource management, task scheduling, cluster governance, disaster recovery planning, cost optimization

Preferred skills

Large-scale distributed system operations, GPU server operations

Responsibilities

Manage and plan GPU/NPU/CPU and storage resources; develop automation tools to improve resource utilization and operational efficiency; oversee multi-region disaster recovery and service deployment; govern cluster machines.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 853,000+ jobs from 20+ sources.