Senior / Staff ML Ops Engineer
Core
Build and evolve training infrastructure for physical AI and autonomous transportation, enabling reliable GPU scheduling, distributed training, and developer tooling for researchers.
Role type
Senior/Staff ML Ops Engineer (Infrastructure)
Builds
Training infrastructure, developer tooling (CLIs, SDKs), data/artifact layers, CI/CD pipelines, and observability systems for ML models.
Domain
Autonomous transportation, Physical AI, Deep Tech
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (GPU scheduling, autoscaling, networking), Python (API/CLI design), AWS (object storage, IAM, GPU compute, IaC), Distributed training (PyTorch DDP/FSDP), Containers, CI/CD, Experiment tracking, Model registry
Preferred skills
Internal developer platforms, Large-scale distributed GPU training, High-throughput sensor data loading, Workflow systems (Argo, Ray, Kubeflow), Simulation infrastructure, Open-source contributions
Technologies
Kubernetes, Helm, Terraform, Pulumi, PyTorch, AWS, Argo Workflows, Ray, Kubeflow, Bazel, Parquet, WebDataset
Responsibilities
Build GPU scheduling and autoscaling on Kubernetes; Design developer-facing CLIs and SDKs; Optimize time-to-first-training and edit-to-signal latency; Evangelize and evaluate ML tooling; Manage dataset versioning and high-throughput loading; Convert one-off scripts into durable, documented libraries; Implement CI/CD for models; Build observability dashboards for the ML stack; Create onboarding guides and support systems; Enforce security and cost governance guardrails.
Seniority
Senior/Staff, hands-on IC with strategic influence