CareerPlanSign in

Senior/Staff Software Engineer, Kubernetes Infrastructure

💼 Full-time💰 $180,000–$180,000🗓 2026-08-14 → 2026-09-25

Core

Design, automate, and deliver high-performance compute environments (bare-metal, VMs, Kubernetes, Slurm) for generative media AI products.

Role type

Senior/Staff Infrastructure Engineer (Kubernetes & GPU)

Builds

Production-grade compute clusters, Linux images, and networking stacks for AI inference and training workloads.

Domain

Generative AI infrastructure, High-Performance Computing (HPC), Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Kubernetes on bare metal, Linux virtualization (KVM/QEMU), NVIDIA GPU stack, Networking (TCP/IP, BGP, VXLAN), Distributed storage, Infrastructure automation (Ansible, Python/Go)

Preferred skills

Slurm, InfiniBand/RoCEv2, SR-IOV/DPDK, OpenStack, KubeVirt, Network automation tools

Technologies

Kubernetes, Slurm, Cilium, Calico, MetalLB, NVIDIA GPU Operator, Ansible, Python, Go, Ceph, Lustre, Weka, KVM, QEMU, InfiniBand, RoCEv2

Responsibilities

Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads; Operate the NVIDIA GPU stack including drivers and device plugins; Design Kubernetes and data-center networking; Build monitoring, alerting, and automated recovery systems; Develop reusable tooling and documentation for infrastructure lifecycle.

Seniority

Senior/Staff, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.