Head of AI Inference & MLOps
Required skills
Significant experience in production AI/LLM inference, MLOps, model serving, or AI infrastructure monetization, Proven experience running or scaling GPU-backed inference systems in production, Strong understanding of modern inference runtimes, serving frameworks, and optimization techniques, Experience with one or more of: vLLM, TensorRT-LLM, SGLang, Ray Serve, Triton Inference Server, Kubernetes-based GPU orchestration, custom routing / scheduler layers, Experience optimizing for real-world production metrics such as throughput, latency, GPU utilization, availability, and cost efficiency, Strong understanding of LLM inference economics, including tradeoffs among model size, quantization, latency, throughput, memory footprint, and customer willingness to pay, Experience building or managing API-based AI platforms or inference products, Ability to translate infrastructure capability into a pricing and product strategy, Experience working with enterprise customers, developer platforms, or AI marketplaces, Strong technical judgment on model selection, infrastructure topology, and commercialization strategy
Preferred skills
Experience monetizing large-scale NVIDIA GPU infrastructure, Experience with rack-scale or cluster-scale inference environments, Background in both technical operations and business strategy, Familiarity with AI inference aggregators, routing platforms, and model marketplaces, Experience designing multi-tenant GPU systems with strong isolation and predictable performance, Experience with advanced observability, token-level metering, cost accounting, and SLA enforcement, Familiarity with reasoning-model workloads, agentic inference, multimodal inference, and future high-density AI factory architectures, Experience supporting OpenAI-compatible APIs and enterprise private deployments
Technologies
vLLM, TensorRT-LLM, SGLang, Ray Serve, Triton Inference Server, Kubernetes-based GPU orchestration, custom routing / scheduler layers, OpenRouter-style aggregation, OpenAI-compatible endpoints, Inference.net
Responsibilities
Build and lead the inference monetization strategy for our first 7MW deployment and expansion to 50MW, Define the technical and commercial operating model for turning GB300 NVL72 racks into revenue-producing assets, Evaluate and implement the model serving stack, scheduling layer, inference engine, observability stack, and API platform, Select and optimize the mix of workloads across: real-time inference reasoning workloads, premium low-latency API traffic, batch / overflow workloads, dedicated enterprise deployments, private/fine-tuned model hosting, Identify the best go-to-market channels for capacity monetization, including direct sales and marketplace/API distribution partners, Develop strategy for integration with platforms such as OpenRouter-style aggregation, OpenAI-compatible endpoints, and other inference distribution channels where appropriate, Own benchmarking methodology based on actual profit and production metrics, not vanity metrics, Drive workload placement decisions based on revenue per rack, revenue per GPU-hour, revenue per MW, latency targets, and customer value, Partner with datacenter engineering, networking, and facilities teams to ensure the physical plant supports the intended software monetization strategy, Build pricing, SLAs, utilization strategy, and customer segmentation framework, Create dashboards and control systems for: utilization, queue health, latency, token throughput, margin by workload, failure rate, realized revenue by cluster / rack / model / customer, Lead decisions around multi-tenant vs single-tenant deployments, reserved vs on-demand capacity, and when to prioritize direct contracts over marketplace traffic, Build and manage the team required to scale this function over time
Seniority
senior
Domain
AI, Machine Learning, Inference, MLOps, Datacenter, Infrastructure, Monetization, Enterprise, API, Cloud