Operations Engineering Manager, Fleet Reliability
Core
Manage the Fleet Reliability Operations team responsible for provisioning, updating, triaging server nodes, and executing processes to configure and validate the server fleet.
Role type
Senior IC manager leading a 24/7 reliability and observability engineering team
Builds
High-volume server fleet lifecycle automation, observability design, and 24/7 engineering support for critical node delivery
Domain
Cloud infrastructure, AI compute, server hardware lifecycle
Deliverable
infrastructure
Required skills
SRE fundamentals, incident management, blameless culture, observability, change management, automation advocacy, talent pipeline development, process documentation, team leadership
Preferred skills
Cross-team process adoption, thought leadership, culture shaping
Technologies
None explicitly stated
Responsibilities
Build and lead a 24/7 team of process-oriented reliability engineers; Socialize and document clear processes for provisioning, validating, and troubleshooting nodes; Advocate for process and automation improvements prioritizing event-driven automated remediation; Provide 24/7 engineering support for high-criticality node delivery; Drive onboarding, documentation, enablement, and performance management programs; Set culture and tone for team communication and collaboration
Seniority
Senior, hands-on IC with leadership scope