CareerPlanSign in

Senior AI Infrastructure Engineer, Kubernetes

San Francisco, California, United States💼 Full-time🗓 2026-09-04 → 2026-09-26

Core

Design and deliver production-grade backend infrastructure for a large-scale GPU cloud platform, focusing on Kubernetes cluster lifecycle, networking, storage, and security for AI workloads.

Role type

Senior IC Kubernetes Infrastructure Engineer (AI/Cloud)

Builds

Production-grade multi-tenant Kubernetes platforms on bare-metal GPU environments

Domain

Cloud Infrastructure / AI Compute / GPU Orchestration

Deliverable

production ML models | infrastructure

Required skills

Kubernetes internals, Infrastructure as Code, Cluster API, GPU device management, Network programming, Storage systems, GitOps/CI/CD, Security engineering, Observability, Mentoring

Preferred skills

Bare-metal provisioning, InfiniBand/RoCE networking, NVIDIA DCGM integration

Technologies

Kubernetes, Cluster API, kubeadm, Metal3, Ceph, NVMe, NVIDIA GPU Operators, SR-IOV, BGP, GitOps, Prometheus, etcd

Responsibilities

Define Kubernetes platform reference architecture and control-plane topology; Build backend services and automation for cluster lifecycle management; Engineer bare-metal deployment workflows using IaC; Design high-performance networking for AI workloads; Implement persistent storage patterns for stateful AI workloads; Integrate NVIDIA GPU operators and device plugins; Establish GitOps and CI/CD patterns; Build platform security controls; Define observability standards and lead failure diagnosis; Set engineering standards and mentor senior engineers

Seniority

Senior, hands-on IC with principal-level responsibilities

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.