CareerPlanSign in

机器学习系统SRE工程师 - Seed Model

上海💼 Full-time🗓 2026-09-28

Core

Maintaining stability of machine learning systems to support large model development, training, and deployment.

Role type

Senior Machine Learning Systems SRE

Builds

Stable GPU infrastructure and resource management for large model training and inference

Domain

AI / Large Language Models / Distributed Systems

Deliverable

infrastructure

Required skills

Linux administration, Go, Python, Shell, Kubernetes, Docker, distributed system resource management, task scheduling, cluster governance, multi-region disaster recovery, cost optimization

Preferred skills

None stated

Technologies

Kubernetes, Docker, Kata, Linux, Go, Python, Shell

Responsibilities

Maintain stability of ML systems supporting model dev/training/deploy; Manage and plan group GPU resources and costs; Govern cluster and service stability to improve resource utilization and operational efficiency; Manage multi-region/multi-datacenter disaster recovery and deployment; Provide operational support and troubleshoot system/business issues

Seniority

Mid-level, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 833,000+ jobs from 20+ sources.