Senior Manager, Production Engineering
Core
Lead and expand the SRE team to ensure reliability, performance, and operational efficiency of the CoreWeave Cloud Platform for AI workloads.
Role type
Senior Manager, Production Engineering (SRE)
Builds
Large-scale, distributed cloud infrastructure for AI training and inference
Domain
Cloud Infrastructure / AI Compute
Deliverable
infrastructure
Required skills
Leadership of distributed 24x7 engineering teams, Incident management and SLO/SLA framework design, Systems engineering for distributed systems, Automation and Infrastructure-as-Code, Cross-functional collaboration
Preferred skills
Platform tooling and internal developer portals, GPU-accelerated workloads management, AI infrastructure environments, Compliance and reliability risk modeling, Bare metal infrastructure environments, DPUs and service mesh architectures
Technologies
Terraform, Kubernetes, Infrastructure-as-Code
Responsibilities
Execute SRE vision and roadmap for distributed cloud infrastructure, Lead and mentor high-performing SRE teams, Champion automation-first practices using AI and IaC, Establish operational excellence best practices, Drive incident management and root cause analysis initiatives, Evolve on-call strategy for 24x7 global platform