CareerPlanSign in

Datacenter Infrastructure Specialist

🌐 Remote💼 Full-time💰 $105,324–$140,432🗓 2026-09-23 → 2026-09-25

Core

Operational linchpin for Runpod's global high-density GPU fleet, bridging hardware partners and internal engineering to ensure uptime for AI/ML workloads.

Role type

Senior IC datacenter infrastructure specialist (HPC/GPU)

Builds

Global physical backbone for AI inference and training

Domain

AI Infrastructure / High-Performance Computing / Datacenter Networking

Deliverable

infrastructure

Required skills

Linux system administration, datacenter networking, GPU stack management, containerization, incident coordination, hardware validation, network troubleshooting

Preferred skills

RDMA/InfiniBand/RoCE, Python/Go/Bash automation, Grafana/Prometheus/Datadog, HPC environment management

Technologies

NVIDIA Software Stack, Docker, RDMA, InfiniBand, RoCE, Grafana, Prometheus, Datadog, Python, Go, Bash

Responsibilities

Validate new hardware against distributed AI/ML specifications; monitor fleet health and audit downtime to protect SLAs; coordinate technical incident communications; support infrastructure partner growth; automate network triage and runbook generation using AI agents.

Seniority

Mid-Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.