AML机器学习系统SRE工程师-火山引擎
Core
Maintain stability of machine learning systems supporting model development, training, and deployment; optimize cluster and service stability, resource utilization, and operational efficiency.
Role type
Senior SRE for ML infrastructure
Builds
Stable ML infrastructure for model training and inference
Domain
Cloud computing / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
Linux administration, C++/Python/Shell scripting, Kubernetes, Docker, distributed system resource management, task scheduling, PyTorch/TensorFlow familiarity, performance optimization, system architecture upgrades
Preferred skills
Deep learning algorithm understanding, algorithm-system joint optimization, business logic abstraction
Technologies
Kubernetes, Docker, PyTorch, TensorFlow, Linux, C++, Python, Shell
Responsibilities
Maintain ML system stability across development, training, and deployment; govern cluster and service stability; optimize resource utilization and operational efficiency; upgrade architecture and optimize data preprocessing/training/inference performance; collaborate with algorithm engineers for joint optimization.
Seniority
Mid-level, hands-on IC