Datacenter Infrastructure Specialist
Core
Operational linchpin for Runpod's global high-density GPU fleet, bridging hardware partners and internal engineering to ensure uptime for AI/ML workloads.
Role type
Senior IC datacenter infrastructure specialist (HPC/GPU)
Builds
Global physical backbone for AI inference and training
Domain
AI Infrastructure / High-Performance Computing / Datacenter Networking
Deliverable
infrastructure
Required skills
Linux system administration, datacenter networking, GPU stack management, containerization, incident coordination, hardware validation, network troubleshooting
Preferred skills
RDMA/InfiniBand/RoCE, Python/Go/Bash automation, Grafana/Prometheus/Datadog, HPC environment management
Technologies
NVIDIA Software Stack, Docker, RDMA, InfiniBand, RoCE, Grafana, Prometheus, Datadog, Python, Go, Bash
Responsibilities
Validate new hardware against distributed AI/ML specifications; monitor fleet health and audit downtime to protect SLAs; coordinate technical incident communications; support infrastructure partner growth; automate network triage and runbook generation using AI agents.
Seniority
Mid-Senior, hands-on IC
