Senior/Staff Software Engineer, Kubernetes Infrastructure
Core
Design, automate, and deliver high-performance compute environments (bare-metal, VMs, Kubernetes, Slurm) for generative media AI products.
Role type
Senior/Staff Infrastructure Engineer (Kubernetes & GPU)
Builds
Production-grade compute clusters, Linux images, and networking stacks for AI inference and training workloads.
Domain
Generative AI infrastructure, High-Performance Computing (HPC), Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes on bare metal, Linux virtualization (KVM/QEMU), NVIDIA GPU stack, Networking (TCP/IP, BGP, VXLAN), Distributed storage, Infrastructure automation (Ansible, Python/Go)
Preferred skills
Slurm, InfiniBand/RoCEv2, SR-IOV/DPDK, OpenStack, KubeVirt, Network automation tools
Technologies
Kubernetes, Slurm, Cilium, Calico, MetalLB, NVIDIA GPU Operator, Ansible, Python, Go, Ceph, Lustre, Weka, KVM, QEMU, InfiniBand, RoCEv2
Responsibilities
Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads; Operate the NVIDIA GPU stack including drivers and device plugins; Design Kubernetes and data-center networking; Build monitoring, alerting, and automated recovery systems; Develop reusable tooling and documentation for infrastructure lifecycle.
Seniority
Senior/Staff, hands-on IC