Staff Engineer, Machine Learning Systems & Reliability - Moveworks
Core
Design and build production systems for the complete ML lifecycle, moving agentic workflows and self-learning models from prototypes to secure, observable, continuously deployable operations.
Role type
Staff Engineer, ML Systems & Reliability
Builds
Production ML platforms, continuous-delivery workflows, self-service automation, and reliable inference/training infrastructure for agentic AI.
Domain
Artificial Intelligence / Machine Learning Systems / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, Go/Java/C++/Rust, distributed systems design, Kubernetes, CI/CD, observability, ML lifecycle management, SRE practices (SLIs/SLOs), capacity planning, technical leadership
Preferred skills
LLMs/agentic automation, cloud infrastructure, infrastructure as code, mentoring
Technologies
Python, Go, Java, C++, Rust, Kubernetes, CI/CD, cloud infrastructure, observability tools
Responsibilities
Design production paths for the ML lifecycle including training, evaluation, serving, and retraining; Implement safe rollout patterns like shadow traffic and canaries; Define and operate SLIs, SLOs, and error budgets; Build self-service platforms to reduce operational toil; Provide technical leadership and mentorship across teams.
Seniority
Staff, hands-on IC with technical leadership