CareerPlanSign in

机器学习系统SRE工程师 - Seed Model

北京💼 Full-time🗓 2026-09-28

Core

Maintaining stability of machine learning systems supporting large model development, training, and deployment.

Role type

Senior Machine Learning Systems SRE

Builds

Stable GPU infrastructure and resource management for large model training and deployment

Domain

AI / Large Language Models / Distributed Systems

Deliverable

infrastructure

Required skills

Linux administration, Go, Python, Shell, Kubernetes, Docker, distributed system resource management, task scheduling, cluster governance, multi-region disaster recovery

Preferred skills

Cost optimization, platform engineering, system stability governance

Responsibilities

Maintain stability of ML systems supporting large model dev/training/deploy; Manage and plan group GPU resources and costs; Govern cluster and service stability to improve resource utilization; Manage multi-region/multi-datacenter disaster recovery and deployment; Provide operational support and troubleshoot system issues

Sourced via bytedance · Listed on CareerPlan, which tracks 833,000+ jobs from 20+ sources.