CareerPlanSign in

Director, Site Operations

Memphis, TN💼 Full-time🗓 2026-09-15 → 2026-09-26

Core

Own node and rack uptime for SpaceXAI's AI supercompute cluster across 5+ sites, ensuring exceptional uptime and customer SLAs.

Role type

Director of Site Operations (Large-scale Data Center/Cluster Reliability)

Builds

AI supercompute cluster infrastructure across multiple sites

Domain

AI/ML infrastructure, High-Performance Computing, Data Center Operations

Deliverable

production ML models | infrastructure

Required skills

Large-scale operations leadership, cluster reliability management, server hardware expertise, vendor management, incident response, SRE leadership, data-driven improvement, cross-functional partnership

Preferred skills

AI/ML compute environment experience, site reliability engineering background, automation tooling familiarity, multi-site scaling experience

Technologies

Jira, Python, Bash

Responsibilities

Own cluster uptime and SLAs across 5+ sites; Lead a 250+ person operations organization; Drive node and rack remediation via physical and command-line intervention; Partner with facilities, network engineering, and tenants; Direct vendor execution for hardware rework; Lead SRE organization for fault mitigation and root cause analysis; Run data-driven improvement initiatives; Command incidents at scale; Scale operations and standardize best practices.

Seniority

Director, hands-on IC with large team leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.