CareerPlanGet AI match score →

Head of AI Inference & MLOps

Austin Area💼 Full-time🗓 2026-03-15 → 2026-07-31

Core

Building a high-density AI datacenter campus focused on real-time inference, reasoning, and high-value AI serving workloads to monetize GPU capacity.

Role type

Head of AI Inference & MLOps (Strategy & Operations)

Builds

Production inference platform for NVIDIA GB300 NVL72 racks, including model serving, routing, and API distribution.

Domain

AI Infrastructure / Datacenter Operations / LLM Inference Economics

Deliverable

production ML models

Required skills

AI/LLM inference strategy, MLOps, model serving and routing, inference economics, GPU cluster optimization, API platform design, multi-tenant architecture, pricing strategy, marketplace integration, observability, team leadership

Preferred skills

NVIDIA GPU infrastructure monetization, rack-scale inference environments, AI inference aggregators, reasoning-model workloads, advanced observability and cost accounting

Technologies

vLLM, TensorRT-LLM, SGLang, Ray Serve, Triton Inference Server, Kubernetes, OpenRouter, Inference.net

Responsibilities

Design and run the inference platform to maximize revenue and utilization; define technical and commercial operating models; select and optimize workload mix; build pricing and SLA frameworks; create dashboards for utilization, latency, and margin; lead team scaling; partner with engineering and facilities teams.

Seniority

Executive, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗