Engineering Manager, Infrastructure Engineering
Core
Lead the MetalDev RAS team to ensure reliability, availability, and serviceability of CoreWeave's bare-metal GPU infrastructure at scale.
Role type
First-line Engineering Manager (Infrastructure/SRE)
Builds
Automation and operational systems for GPU server fleets
Domain
Cloud infrastructure / Bare-metal hardware / AI compute
Deliverable
production ML models | infrastructure
Required skills
Engineering management, Site Reliability Engineering (SRE), Cloud operations, Kubernetes (K8S), Go programming, Incident management, Observability, SLA/SLO definition
Preferred skills
Bare-metal/hardware infrastructure management, Redfish/BMC technologies, High-growth scaling
Technologies
Prometheus, Grafana, Redfish, BMC, CI/CD pipelines, Kubernetes
Responsibilities
Hire, coach, and grow a team of infrastructure/SRE engineers; Own RAS metrics and drive operational excellence; Lead incident response and RCA processes; Collaborate cross-functionally on platform reliability; Engage with upstream communities to guide technical standards
Seniority
Manager, first-line leadership