Senior Kubernetes Engineer
Core
Design, implement, and optimize GPU-accelerated container platforms at scale for AI/ML, HPC, and LLM training workloads in hybrid or on-prem environments.
Role type
Senior IC Kubernetes Engineer (GPU infrastructure)
Builds
GPU-accelerated Kubernetes clusters and custom operators for high-performance workloads
Domain
High-performance computing (HPC) and cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals, GPU scheduling, NVIDIA ecosystem (Operator, device plugins, MIG, DCGM), Go or Python, custom operators/controllers, scheduler extensions, Helm, Kustomize, GitOps, Terraform, CNI plugins, monitoring (Prometheus, DCGM Exporter)
Preferred skills
None stated
Technologies
Kubernetes, NVIDIA GPU Operator, NVIDIA Device Plugin, kube-scheduler plugins, Slurm, Volcano, Prometheus, Grafana, DCGM Exporter, OpenTelemetry, OPA, Gatekeeper, ArgoCD, FluxCD, Terraform, Helm, Kustomize
Responsibilities
Architecting and operating Kubernetes clusters optimised for GPU workloads; Developing, deploying and maintaining custom Kubernetes operators and controllers; Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features; Optimising GPU utilisation and job placement through scheduler extensions; Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance; Driving observability and telemetry integrations; Implementing secure multi-user and multi-namespace GPU isolation; Maintaining CI/CD pipelines for Kubernetes infrastructure; Contributing to infrastructure-as-code; Participating in performance tuning, incident response and production readiness reviews
Seniority
Senior, hands-on IC