CareerPlanSign in

Production Engineer, IaaS

San Francisco, CA💼 Full-time🗓 2026-07-05 → 2026-09-26

Core

Build and operate the observability platform, control plane, and fleet state management for a hyperscale GPU/IaaS infrastructure supporting frontier AI compute.

Role type

Senior IC Production Engineer (IaaS/Infrastructure)

Builds

Production-grade observability pipelines, unified control plane APIs, and fleet state integration for tens of thousands of GPUs.

Domain

AI Infrastructure / Hyperscale Data Centers / IaaS

Deliverable

production ML models | infrastructure

Required skills

API design and versioning, distributed systems, observability stack engineering, incident management, AI tooling fluency (LLM APIs, agentic frameworks), hardware telemetry integration.

Preferred skills

Go, Python, Postgres, Kubernetes, time-series databases (Prometheus, Thanos, VictoriaMetrics), workflow orchestration (Temporal, Cadence), BMC/Redfish.

Technologies

Kubernetes, Prometheus, Thanos, VictoriaMetrics, Temporal, Cadence, Go, Python, Postgres, LLM APIs, MCP servers.

Responsibilities

Own the observability platform data pipelines and healthcheck frameworks; Define and build the unified API surface for infrastructure management; Build the production control plane for machine management and command execution; Maintain fleet state as the source of truth across provisioning and operations; Integrate new hardware generations (ZTP, DHCP, DNS) into the platform.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.