Executive Director, AI Infrastructure & Platform Engineering
Core
Building and operating a frontier-class on-premises AI compute platform (GPU clusters, network fabric, storage) for CVS Health's Enterprise AI Factory.
Role type
Senior Engineering Leadership (Executive Director)
Builds
Physical AI infrastructure, bare-metal Kubernetes/OpenShift clusters, high-performance RoCE v2 network fabric, and 24/7 SRE operations.
Domain
Healthcare technology, Data Center Infrastructure, High-Performance Computing (HPC)
Deliverable
Infrastructure
Required skills
Executive leadership, physical data center operations, bare-metal Kubernetes/OpenShift, high-speed cluster fabrics (RoCE v2, InfiniBand), SRE practices, FinOps, vendor management, HIPAA compliance
Preferred skills
NVIDIA Blackwell/HGX/DGX systems, VAST distributed NVMe storage, GPU cluster operations (32+ GPUs), chaos engineering, innovation program management
Technologies
NVIDIA Blackwell, RoCE v2, Kubernetes, OpenShift, Cisco UCS, VAST
Responsibilities
Define long-range strategy for AI infrastructure, recruit and develop engineering teams, own physical layer operations (GPU, storage, network), enforce operational baselines and SLOs, manage FinOps and vendor relationships, lead organizational transition to permanent operations.
Seniority
Executive Director, hands-on IC leadership