CareerPlanSign in

Machine Learning Ops Engineer, Global SRE

San Jose, United States of America💼 Full-time🗓 2026-09-28

Core

Ensuring stability and efficiency of machine learning systems across the full lifecycle from data preparation to serving, including offline training and online inference.

Role type

Senior IC MLOps/SRE engineer

Builds

Stable online ML serving systems, offline training pipelines, and GPU-based model training infrastructure

Domain

Machine Learning Operations / Site Reliability Engineering

Deliverable

infrastructure

Required skills

Linux systems administration, networking, storage management, Python, Go, C, C++, or Java, production operations troubleshooting, SLO definition

Preferred skills

SRE experience for ML systems, SRE experience for ads/recommendation/search systems

Technologies

Linux, Python, Go, C, C++, Java, GPU

Responsibilities

Set and maintain SLOs for online ML serving systems, maintain stability of offline ML training tasks, roll out GPU model training in Non-China regions, ensure stability of AIGC related ML tasks, manage and plan ML resources including cost, budget, and efficiency

Seniority

Senior, hands-on IC

Sourced via tiktok · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.