CareerPlanSign in

Sr MLOps Engineer

Sunnyvale, CA, us💼 Full-time🗓 2026-08-06 → 2026-09-25

Core

Design, build, and maintain infrastructure and tools for the full machine learning lifecycle, ensuring seamless integration of ML models into production systems.

Role type

Senior MLOps Engineer

Builds

Production-grade ML infrastructure, orchestration tooling, and artifact/dataset storage solutions

Domain

Healthcare technology / Robotic surgery / Cloud infrastructure

Deliverable

infrastructure

Required skills

Kubernetes production operations, Python/Bash scripting, Infrastructure-as-Code (Terraform/Helm/Ansible), distributed storage systems, CI/CD pipelines, Linux system administration, GPU hardware management

Preferred skills

ML orchestration frameworks (Metaflow/MLflow/Kubeflow), large-scale infrastructure migration leadership, regulated industry experience

Technologies

Kubernetes, Metaflow, S3, MinIO, NVIDIA GPUs (B200/L40S/A6000/V100), CUDA, GitLab CI, ArgoCD

Responsibilities

Bootstrap and maintain production Kubernetes clusters with CNI and storage integration; Deploy and configure ML orchestration tooling and artifact storage; Validate GPU node health and configuration across heterogeneous hardware; Design and execute team migration playbooks for workflow and dataset porting; Write and maintain runbooks, architecture documentation, and disaster recovery procedures; Participate in on-call rotation and incident response; Collaborate with IT/Security on identity integration and compliance

Seniority

Senior, hands-on IC

Sourced via smartrecruiters · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.