Machine Learning Developer
Core
End-to-end operations for production ML models, including monitoring, retraining, and incident response for agentic/LLM systems.
Role type
Senior Machine Learning Operations Engineer
Builds
Production ML models and operational workflows for client projects
Domain
IT Services / Cloud Infrastructure (Google Cloud)
Deliverable
production ML models
Required skills
Model auditing, performance tuning, data pipeline management, cloud environment management, incident response, drift detection, SLA reporting, model onboarding, automation scripting
Preferred skills
Experience with Vertex AI, structured intake processes, technical debt reduction
Technologies
Google Cloud, Vertex AI, pipelines, monitoring tools
Responsibilities
Perform deep dive audits of ML models and create Model Runbooks; Own end to end model operations including backlog execution and performance tuning; Manage retraining cycles, data pipelines, cloud environments, and deployment workflows; Monitor model health, detect drift, and respond to incidents; Deliver monthly SLA dashboards and performance reports; Onboard new ML models using a structured intake and evaluation process; Support agentic/LLM systems with incident response, bug triage, and rollbacks; Automate operational tasks and reduce technical debt
Seniority
Senior, hands-on IC