CareerPlanSign in

Inference Engineer

San Francisco, CA💼 Full-time🗓 2026-09-09 → 2026-09-26

Core

Build end-to-end inference capabilities on a unified control plane (Forge) to deploy and serve AI models across distributed, heterogeneous hardware clusters.

Role type

Senior Inference Engineer (Infrastructure)

Builds

Production-ready model serving infrastructure, monitoring, gateways, and endpoints for a global GPU marketplace.

Domain

AI Infrastructure / GPU Cloud / Distributed Systems

Deliverable

production ML models

Required skills

Kubernetes cluster operations, inference frameworks evaluation, NVIDIA Dynamo architecture, KV-cache orchestration, autoscaling, production monitoring setup, end-to-end product building

Preferred skills

Model optimization (quantization, batching, kernel tuning), RDMA networking, heterogeneous accelerator deployment, customer debugging support

Technologies

Kubernetes, NVIDIA Dynamo, Forge, NeoCloud

Responsibilities

Deploy and serve models on heterogeneous hardware clusters, evaluate and select inference frameworks, build monitoring and gateways for production readiness, optimize inference performance and manage autoscaling, orchestrate KV-cache, debug customer inference issues

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.