Infrastructure Operations Engineer
Core
Design, deploy, and operate OpenStack and Kubernetes environments for GPU workloads, ensuring performance, scalability, and resilience.
Role type
Senior Infrastructure Operations Engineer (via careerplan.io/jobs/4975430101-infrastructure-operations-engineer-at-nexgen-cloud)
Builds
On-demand and private GPU infrastructure for AI researchers and enterprises
Domain
AI Cloud Infrastructure / HPC
Deliverable
infrastructure
Required skills
Linux systems administration, hardware assembly and racking, data center operations, networking, storage systems
Preferred skills
GPU hardware installation and configuration, OpenStack, Kubernetes, infrastructure automation, CI/CD, Git workflows, HPC environments
Technologies
OpenStack, Kubernetes, NVIDIA tooling
Responsibilities
Own the design, deployment, and operation of OpenStack and Kubernetes environments; Build and improve infrastructure using infrastructure-as-code and GitOps practices; Optimise GPU workload scheduling and implement monitoring, logging, and alerting; Lead incident response and drive continuous improvement of reliability; Maintain strong security controls across infrastructure and container layers; Work closely with Platform, DevOps, AI, Product, and Support teams
Seniority
Senior, hands-on IC