CareerPlanSign in

Senior Machine Learning Platform Ops

USA💼 Full-time🗓 2026-09-24 → 2026-09-26

Core

Build and maintain scalable ML pipelines, containerized training environments, and observability systems for production ML and LLM features.

Role type

Senior Machine Learning Platform Engineer (Ops/DevOps)

Builds

Production ML pipelines, containerized model-training environments, LLM serving infrastructure, and internal platform tooling.

Domain

Cloud Infrastructure, Machine Learning Operations, Generative AI

Deliverable

production ML models | infrastructure

Required skills

Python, SQL, Kubernetes, Docker, Terraform, Git, CI/CD, Workflow Orchestration (Airflow/Kubeflow/Dagster), Cloud Platforms (GCP/AWS), ML Lifecycle Management, Observability, GPU/Spot Autoscaling

Preferred skills

LLM serving, Vector databases, Agentic AI SDLC practices, Mentorship

Technologies

Python, SQL, Airflow, Kubeflow, Dagster, GCP, AWS, Kubernetes, Docker, Terraform, Git

Responsibilities

Build and maintain ML pipelines for training, evaluation, and deployment; Create reproducible, containerized model-training environments; Define observability and alerting for ML systems; Design and scale batch and streaming data-ingestion flows; Develop internal Python libraries and platform tooling; Explore and productionize LLM-based features; Mentor peers in reliability and testing.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.