Manager, Site Operations
Core
Oversee data center technicians and infrastructure to ensure 99.999% uptime for AI compute systems, managing power, cooling, networking, and hardware deployments.
Role type
Manager, Site Operations (Data Center)
Builds
Reliable AI compute infrastructure and sustainable data center operations
Domain
Data Center Operations / AI Infrastructure
Deliverable
infrastructure
Required skills
Data center operations management, team leadership, server hardware expertise, incident resolution, inventory management, vendor coordination, sustainability optimization, metrics tracking, emergency response, process optimization
Preferred skills
AI/ML compute environment experience, Jira workflow management, technical communication, scripting (Python/Bash), vendor partnership, scaling operations
Technologies
Jira, Python, Bash
Responsibilities
Manage power, cooling, networking, and hardware deployments; Lead and develop a team of Data Center Operations Technicians; Take charge of hardware lifecycles, incident resolution, and inventory management; Coordinate between technicians, AI specialists, and external vendors; Champion energy-efficient practices and sustainability efforts; Track and report key metrics like uptime and power efficiency; Lead the team through urgent situations; Build and refine processes for preventative maintenance and ticket workflows; Work with leadership to standardize best practices across sites
Seniority
Manager, hands-on IC with team leadership