Director, Site Operations
Core
Own node and rack uptime for SpaceXAI's AI supercompute cluster across 5+ sites, ensuring exceptional uptime and customer SLAs.
Role type
Director of Site Operations (Large-scale Data Center/Cluster Reliability)
Builds
AI supercompute cluster infrastructure across multiple sites
Domain
AI/ML infrastructure, High-Performance Computing, Data Center Operations
Deliverable
production ML models | infrastructure
Required skills
Large-scale operations leadership, cluster reliability management, server hardware expertise, vendor management, incident response, SRE leadership, data-driven improvement, cross-functional partnership
Preferred skills
AI/ML compute environment experience, site reliability engineering background, automation tooling familiarity, multi-site scaling experience
Technologies
Jira, Python, Bash
Responsibilities
Own cluster uptime and SLAs across 5+ sites; Lead a 250+ person operations organization; Drive node and rack remediation via physical and command-line intervention; Partner with facilities, network engineering, and tenants; Direct vendor execution for hardware rework; Lead SRE organization for fault mitigation and root cause analysis; Run data-driven improvement initiatives; Command incidents at scale; Scale operations and standardize best practices.
Seniority
Director, hands-on IC with large team leadership