Head of AI Inference & MLOps
Core
Building a high-density AI datacenter campus focused on real-time inference, reasoning, and high-value AI serving workloads to monetize GPU capacity.
Role type
Head of AI Inference & MLOps (Strategy & Operations)
Builds
Production inference platform for NVIDIA GB300 NVL72 racks, including model serving, routing, and API distribution.
Domain
AI Infrastructure / Datacenter Operations / LLM Inference Economics
Deliverable
production ML models
Required skills
AI/LLM inference strategy, MLOps, model serving and routing, inference economics, GPU cluster optimization, API platform design, multi-tenant architecture, pricing strategy, marketplace integration, observability, team leadership
Preferred skills
NVIDIA GPU infrastructure monetization, rack-scale inference environments, AI inference aggregators, reasoning-model workloads, advanced observability and cost accounting
Technologies
vLLM, TensorRT-LLM, SGLang, Ray Serve, Triton Inference Server, Kubernetes, OpenRouter, Inference.net
Responsibilities
Design and run the inference platform to maximize revenue and utilization; define technical and commercial operating models; select and optimize workload mix; build pricing and SLA frameworks; create dashboards for utilization, latency, and margin; lead team scaling; partner with engineering and facilities teams.
Seniority
Executive, hands-on IC