Senior AI Infrastructure Engineer
Core
Build and optimize the engine behind large-scale AI capacity by establishing a high-performance distributed system for GPU compute.
Role type
Senior hands-on AI Infrastructure Engineer
Builds
Distributed GPU clusters, high-performance networks, and AI storage platforms as a managed service
Domain
AI Infrastructure / HPC / Cloud Native
Deliverable
production ML models | infrastructure
Required skills
Linux, Kubernetes, Infrastructure as Code, GPU cluster management, high-performance networking, system debugging, capacity planning, observability
Preferred skills
Nvidia DGX/HGX, CUDA, InfiniBand, RDMA, Slurm, bare metal automation, confidential computing
Technologies
Linux, Kubernetes, Ansible, Terraform, Nvidia, InfiniBand, Slurm
Responsibilities
Build and configure Nvidia-based GPU clusters for training and inference; Integrate compute, network, and storage; Implement Linux, Kubernetes, and workload scheduling; Automate provisioning and upgrades with IaC; Perform performance testing and root cause analysis; Build observability and logging; Document implementations and runbooks
Seniority
Senior, hands-on IC