Machine Learning Ops Engineer, Global SRE
Core
Ensuring stability and efficiency of machine learning systems across the full lifecycle from data preparation to serving, including offline training and online inference.
Role type
Senior IC MLOps/SRE engineer
Builds
Stable online ML serving systems, offline training pipelines, and GPU-based model training infrastructure
Domain
Machine Learning Operations / Site Reliability Engineering
Deliverable
infrastructure
Required skills
Linux systems administration, networking, storage management, Python, Go, C, C++, or Java, production operations troubleshooting, SLO definition
Preferred skills
SRE experience for ML systems, SRE experience for ads/recommendation/search systems
Technologies
Linux, Python, Go, C, C++, Java, GPU
Responsibilities
Set and maintain SLOs for online ML serving systems, maintain stability of offline ML training tasks, roll out GPU model training in Non-China regions, ensure stability of AIGC related ML tasks, manage and plan ML resources including cost, budget, and efficiency
Seniority
Senior, hands-on IC


