CareerPlanSign in

AML机器学习平台SRE工程师-Data

深圳💼 Full-time🗓 2026-09-28

Core

Ensure stability of machine learning systems supporting model development, training, and deployment; manage GPU/NPU/CPU and storage resources; handle multi-region disaster recovery and cluster governance.

Role type

Senior Site Reliability Engineer (ML Infrastructure)

Builds

ML training and inference platforms, resource management systems, automation tools

Domain

Cloud infrastructure, Machine Learning, Distributed Systems

Deliverable

infrastructure

Required skills

Linux administration, Go, Python, Shell scripting, Kubernetes, Distributed system operations, Resource scheduling, Cluster governance, Disaster recovery planning, Cost optimization

Preferred skills

Large-scale distributed system operations, GPU server management

Technologies

Kubernetes, Linux, Go, Python, Shell, GPU/NPU/CPU, Storage systems

Responsibilities

Manage and plan GPU/NPU/CPU and storage resources including cost and budget; Implement multi-region disaster recovery and service deployment; Develop automation tools to improve resource utilization and operational efficiency; Maintain cluster machine governance.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.