CareerPlanSign in

AML机器学习系统SRE工程师

杭州💼 Full-time🗓 2026-09-28

Core

Maintaining stability of machine learning systems, supporting model development, training, and deployment, managing resources, and ensuring disaster recovery across multi-region clusters.

Role type

Senior SRE for Applied Machine Learning systems

Builds

High-performance, scalable ML system architecture and end-to-end ML services

Domain

Machine Learning Infrastructure / Distributed Systems

Deliverable

infrastructure

Required skills

Linux administration, Go, Python, Shell, Kubernetes, Docker, Kata, resource management, cost optimization, disaster recovery, cluster governance

Preferred skills

Large-scale distributed system operations, GPU server operations

Technologies

Kubernetes, Docker, Kata, Linux, Go, Python, Shell

Responsibilities

Maintain stability of ML systems across model development, training, and deployment; Manage and plan GPU/CPU resources and storage; Handle multi-region disaster recovery and cluster governance; Improve system stability, resource utilization, and operational efficiency.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.