CareerPlanSign in

Senior / Staff ML Ops Engineer

💼 Full-time🗓 2026-09-23 → 2026-09-25

Core

Build and evolve training infrastructure for physical AI and autonomous transportation, enabling reliable GPU scheduling, distributed training, and developer tooling for researchers.

Role type

Senior/Staff ML Ops Engineer (Infrastructure)

Builds

Training infrastructure, developer tooling (CLIs, SDKs), data/artifact layers, CI/CD pipelines, and observability systems for ML models.

Domain

Autonomous transportation, Physical AI, Deep Tech

Deliverable

production ML models | infrastructure

Required skills

Kubernetes (GPU scheduling, autoscaling, networking), Python (API/CLI design), AWS (object storage, IAM, GPU compute, IaC), Distributed training (PyTorch DDP/FSDP), Containers, CI/CD, Experiment tracking, Model registry

Preferred skills

Internal developer platforms, Large-scale distributed GPU training, High-throughput sensor data loading, Workflow systems (Argo, Ray, Kubeflow), Simulation infrastructure, Open-source contributions

Technologies

Kubernetes, Helm, Terraform, Pulumi, PyTorch, AWS, Argo Workflows, Ray, Kubeflow, Bazel, Parquet, WebDataset

Responsibilities

Build GPU scheduling and autoscaling on Kubernetes; Design developer-facing CLIs and SDKs; Optimize time-to-first-training and edit-to-signal latency; Evangelize and evaluate ML tooling; Manage dataset versioning and high-throughput loading; Convert one-off scripts into durable, documented libraries; Implement CI/CD for models; Build observability dashboards for the ML stack; Create onboarding guides and support systems; Enforce security and cost governance guardrails.

Seniority

Senior/Staff, hands-on IC with strategic influence

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.