Software Engineer, Infrastructure
Core
Build software and processes to manage a large fleet of GPU servers, ensuring high availability, health monitoring, and automated recovery for AI workloads.
Role type
Senior Infrastructure Engineer (GPU Fleet Management)
Builds
Production-grade server management tooling, fleet tracking systems, and automated recovery processes for thousands of GPU nodes.
Domain
AI Infrastructure / High-Performance Computing / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Python, Linux systems administration, Infrastructure-as-Code (Ansible, Terraform), Storage technologies (NVMe, NFS, LVM), GPU diagnostics, Network configuration, OS hardening
Preferred skills
AI-driven automation, Distributed storage systems, NVIDIA/AMD GPU stack optimization, Compliance frameworks (SOC 2, ISO 27001)
Technologies
Python, Ansible, Terraform, NVMe, NFS, Lustre, GPFS, NVIDIA drivers, CUDA, SELinux, AppArmor, BGP, VXLAN, InfiniBand
Responsibilities
Build and maintain Python fleet tracking systems for server lifecycle management; Develop automated provisioning, health checks, and recovery tooling; Create metrics dashboards and alerting for hardware health; Tune Linux systems and GPU drivers for AI workloads; Implement OS-level security hardening and compliance automation; Manage distributed and local storage systems for model weights and checkpoints.
Seniority
Senior, hands-on IC