Manager, Production Engineering
Core
Founding Engineering Manager building the Cloud Availability Platform (CAPE) reliability spine for Crusoe's vertically integrated AI neocloud, bridging physical infrastructure with software-defined operations.
Role type
Senior IC Engineering Manager (Production/SRE)
Builds
Cloud Availability Platform (CAPE) for Crusoe's AI neocloud
Domain
AI Infrastructure / Data Center Operations / Energy
Deliverable
production ML models | infrastructure
Required skills
Infrastructure/SRE/Production Engineering leadership, Hands-on systems programming (Go/Python/C++), Distributed systems expertise, Linux internals, Container orchestration, On-call management, SLI/SLO definition, Automation/Tooling design
Preferred skills
AI infrastructure/GPU cluster experience, High-performance fabric (InfiniBand/RoCEv2), Hardware internals (BMC/firmware), Accelerator failure mode knowledge, Large-scale fleet scaling
Technologies
Go, Python, C++, Temporal, Linux, Kubernetes, InfiniBand, RoCEv2
Responsibilities
Recruit and mentor a local Production Engineering team in Tel Aviv, Own global on-call rotation and blameless post-mortem culture, Drive software-defined operations and alert reduction, Act as Production Gatekeeper for cross-functional readiness reviews, Ensure team bandwidth is allocated to strategic automation, Convert physical interventions into software-defined auto-remediation