Senior Machine Learning Platform Ops
Core
Build and maintain scalable ML pipelines, containerized training environments, and observability systems for production ML and LLM features.
Role type
Senior Machine Learning Platform Engineer (Ops/DevOps)
Builds
Production ML pipelines, containerized model-training environments, LLM serving infrastructure, and internal platform tooling.
Domain
Cloud Infrastructure, Machine Learning Operations, Generative AI
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, Kubernetes, Docker, Terraform, Git, CI/CD, Workflow Orchestration (Airflow/Kubeflow/Dagster), Cloud Platforms (GCP/AWS), ML Lifecycle Management, Observability, GPU/Spot Autoscaling
Preferred skills
LLM serving, Vector databases, Agentic AI SDLC practices, Mentorship
Technologies
Python, SQL, Airflow, Kubeflow, Dagster, GCP, AWS, Kubernetes, Docker, Terraform, Git
Responsibilities
Build and maintain ML pipelines for training, evaluation, and deployment; Create reproducible, containerized model-training environments; Define observability and alerting for ML systems; Design and scale batch and streaming data-ingestion flows; Develop internal Python libraries and platform tooling; Explore and productionize LLM-based features; Mentor peers in reliability and testing.
Seniority
Senior, hands-on IC with mentorship responsibilities