Senior AI Infrastructure Engineer, Kubernetes
Core
Design and deliver production-grade backend infrastructure for a large-scale GPU cloud platform, focusing on Kubernetes cluster lifecycle, networking, storage, and security for AI workloads.
Role type
Senior IC Kubernetes Infrastructure Engineer (AI/Cloud)
Builds
Production-grade multi-tenant Kubernetes platforms on bare-metal GPU environments
Domain
Cloud Infrastructure / AI Compute / GPU Orchestration
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals, Infrastructure as Code, Cluster API, GPU device management, Network programming, Storage systems, GitOps/CI/CD, Security engineering, Observability, Mentoring
Preferred skills
Bare-metal provisioning, InfiniBand/RoCE networking, NVIDIA DCGM integration
Technologies
Kubernetes, Cluster API, kubeadm, Metal3, Ceph, NVMe, NVIDIA GPU Operators, SR-IOV, BGP, GitOps, Prometheus, etcd
Responsibilities
Define Kubernetes platform reference architecture and control-plane topology; Build backend services and automation for cluster lifecycle management; Engineer bare-metal deployment workflows using IaC; Design high-performance networking for AI workloads; Implement persistent storage patterns for stateful AI workloads; Integrate NVIDIA GPU operators and device plugins; Establish GitOps and CI/CD patterns; Build platform security controls; Define observability standards and lead failure diagnosis; Set engineering standards and mentor senior engineers
Seniority
Senior, hands-on IC with principal-level responsibilities