CareerPlanGet AI match score →

Operations Engineering Manager, Fleet Reliability

Dallas, TX💼 Full-time💰 $143,000–$191,000🗓 2026-07-17 → 2026-07-31

Core

Manage the Fleet Reliability Operations team responsible for provisioning, updating, triaging server nodes, and executing processes to configure and validate the server fleet.

Role type

Senior IC manager leading a 24/7 reliability and observability engineering team

Builds

High-volume server fleet lifecycle automation, observability design, and 24/7 engineering support for critical node delivery

Domain

Cloud infrastructure, AI compute, server hardware lifecycle

Deliverable

infrastructure

Required skills

SRE fundamentals, incident management, blameless culture, observability, change management, automation advocacy, talent pipeline development, process documentation, team leadership

Preferred skills

Cross-team process adoption, thought leadership, culture shaping

Technologies

None explicitly stated

Responsibilities

Build and lead a 24/7 team of process-oriented reliability engineers; Socialize and document clear processes for provisioning, validating, and troubleshooting nodes; Advocate for process and automation improvements prioritizing event-driven automated remediation; Provide 24/7 engineering support for high-criticality node delivery; Drive onboarding, documentation, enablement, and performance management programs; Set culture and tone for team communication and collaboration

Seniority

Senior, hands-on IC with leadership scope

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗