CareerPlanSign in

Staff Engineer - ML Infra / MLOps

United States - Palo Alto, CA💼 Full-time💰 $218,000–$218,000🗓 2026-07-23 → 2026-09-26

Core

Design and operate production-grade ML infrastructure platforms enabling distributed training, feature stores, and high-throughput inference serving for a scaling retail AI team.

Role type

Staff ML Infrastructure Engineer (MLOps)

Builds

End-to-end ML platform including training pipelines, serving infrastructure, and developer tooling for data scientists.

Domain

Retail / Cloud-native ML Infrastructure

Deliverable

production ML models

Required skills

Distributed training pipelines, Feature stores, High-throughput inference serving, GPU utilization optimization, Cloud cost control, Kubernetes (EKS), Infrastructure as Code (Terraform/Pulumi), ML frameworks (PyTorch, TensorFlow, Kubeflow, SageMaker), Data pipelines (Spark, Flink, Kafka), CI/CD for ML, Model versioning, Experiment tracking, Deployment strategies (blue-green, canary), Root-cause analysis, On-call discipline

Preferred skills

Startup hustle, Handling ambiguity, Rapid experimentation

Technologies

AWS, Kubernetes, Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka

Responsibilities

Architect end-to-end ML platform design for scalability and extensibility; Build developer experience tooling to reduce friction from idea to production; Set engineering standards for CI/CD, IaC, and deployment strategies; Lead technical evaluation of core platform components; Optimize GPU utilization and cloud costs; Ensure production reliability and automated recovery; Mentor junior and mid-level engineers; Lead root-cause analyses for production failures.

Seniority

Staff, hands-on IC with mentorship

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.