CareerPlanSign in

Software Engineer, Infrastructure

San Francisco💼 Full-time💰 $180,000–$250,000🗓 2026-09-22 → 2026-09-25

Core

Build software and processes to manage a large fleet of GPU servers, ensuring high availability, health monitoring, and automated recovery for AI workloads.

Role type

Senior Infrastructure Engineer (GPU Fleet Management)

Builds

Production-grade server management tooling, fleet tracking systems, and automated recovery processes for thousands of GPU nodes.

Domain

AI Infrastructure / High-Performance Computing / Cloud Infrastructure

Deliverable

infrastructure

Required skills

Python, Linux systems administration, Infrastructure-as-Code (Ansible, Terraform), Storage technologies (NVMe, NFS, LVM), GPU diagnostics, Network configuration, OS hardening

Preferred skills

AI-driven automation, Distributed storage systems, NVIDIA/AMD GPU stack optimization, Compliance frameworks (SOC 2, ISO 27001)

Technologies

Python, Ansible, Terraform, NVMe, NFS, Lustre, GPFS, NVIDIA drivers, CUDA, SELinux, AppArmor, BGP, VXLAN, InfiniBand

Responsibilities

Build and maintain Python fleet tracking systems for server lifecycle management; Develop automated provisioning, health checks, and recovery tooling; Create metrics dashboards and alerting for hardware health; Tune Linux systems and GPU drivers for AI workloads; Implement OS-level security hardening and compliance automation; Manage distributed and local storage systems for model weights and checkpoints.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.