CareerPlanSign in

AML机器学习系统SRE工程师-火山引擎

北京💼 Full-time🗓 2026-09-28

Core

Maintain stability of machine learning systems supporting model development, training, and deployment; optimize cluster and service stability, resource utilization, and operational efficiency.

Role type

Senior SRE for ML infrastructure

Builds

Stable ML infrastructure for model training and inference

Domain

Cloud computing / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

Linux administration, C++/Python/Shell scripting, Kubernetes, Docker, distributed system resource management, task scheduling, PyTorch/TensorFlow familiarity, performance optimization, system architecture upgrades

Preferred skills

Deep learning algorithm understanding, algorithm-system joint optimization, business logic abstraction

Technologies

Kubernetes, Docker, PyTorch, TensorFlow, Linux, C++, Python, Shell

Responsibilities

Maintain ML system stability across development, training, and deployment; govern cluster and service stability; optimize resource utilization and operational efficiency; upgrade architecture and optimize data preprocessing/training/inference performance; collaborate with algorithm engineers for joint optimization.

Seniority

Mid-level, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.