CareerPlanGet AI match score →

Senior Cloud Infrastructure Engineer

Santa Clara, CA💼 Full-time💰 $180,000–$180,000🗓 2026-07-01 → 2026-07-31

Core

Architect and manage large-scale compute and data infrastructure powering autonomous driving stacks, ensuring multi-GPU clusters and distributed training frameworks are scalable and resilient.

Role type

Senior Cloud Infrastructure Engineer (MLOps)

Builds

High-performance GPU/TPU clusters, data pipelines for sensor processing, and automated model lifecycle workflows for autonomous vehicle simulation and training.

Domain

Autonomous driving / AI infrastructure / Cloud computing

Deliverable

production ML models

Required skills

Kubernetes (K8s) mastery, Infrastructure as Code (Terraform, Helm), Distributed systems (Ray, PyTorch Distributed), Data engineering (Apache Airflow, Kafka, Spark), Observability (Prometheus, Grafana, OpenTelemetry), Python, Bash scripting, IAM/RBAC

Preferred skills

Distributed training frameworks (FSDP, DeepSpeed), AI Agent orchestration (LangGraph, AutoGen), Advanced networking protocols (InfiniBand, RoCE v2), Model Context Protocol (MCP)

Technologies

Kubernetes, NVIDIA GPU Operator, Terraform, Helm, Apache Airflow, Kafka, Spark, ArgoCD, Gitlab CI/CD, MLFlow, Triton Inference Server, Ray Serve, ONNX Runtime, PyTorch, TorchElastic, Horovod, Prometheus, Grafana, OpenTelemetry, LangGraph, CrewAI

Responsibilities

Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads; Implement and optimize Kubernetes-native GPU scheduling; Build large-scale data pipelines to process raw sensor data; Develop agent-driven CI/CD workflows for infrastructure and model artifacts; Design and maintain MLFlow and feature store integrations for model tracking; Optimize low-level communication (NCCL, InfiniBand) for large-scale distributed training.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗