Director of Engineering, Flex Compute
Core
Architecting and building the Flexible Compute system to manage data-center power draw, grid integration, and dynamic power management for AI workloads.
Role type
Director of Engineering, Flex Compute (0-1 build)
Builds
Utility-validated pilot for flexible load, dynamic power management system, GPU power estimation engine, and safety-critical control plane.
Domain
AI Infrastructure / Energy Grid / Data Center Operations
Deliverable
production ML models | infrastructure
Required skills
Distributed control planes, fleet automation, safety-critical systems, data-center power systems, energy markets/grid programs, build-vs-leverage judgment, hiring senior engineers
Preferred skills
GPU cluster operations, energy sector background, GPU power modeling, checkpoint/preemption
Technologies
Kubernetes, Slurm, Temporal, BESS, UPS, switchgear
Responsibilities
Define architecture and ship first production system, own curtailment orchestration and decision layer, land utility-validated pilot, design for safety and graceful ride-through of grid events, partner with Data Center Engineering and Energy teams, integrate with cloud control plane, hire and lead team from scratch
Seniority
Director, hands-on IC with team leadership