Production Engineer, IaaS
Core
Build and operate the observability platform, control plane, and fleet state management for a hyperscale GPU/IaaS infrastructure supporting frontier AI compute.
Role type
Senior IC Production Engineer (IaaS/Infrastructure)
Builds
Production-grade observability pipelines, unified control plane APIs, and fleet state integration for tens of thousands of GPUs.
Domain
AI Infrastructure / Hyperscale Data Centers / IaaS
Deliverable
production ML models | infrastructure
Required skills
API design and versioning, distributed systems, observability stack engineering, incident management, AI tooling fluency (LLM APIs, agentic frameworks), hardware telemetry integration.
Preferred skills
Go, Python, Postgres, Kubernetes, time-series databases (Prometheus, Thanos, VictoriaMetrics), workflow orchestration (Temporal, Cadence), BMC/Redfish.
Technologies
Kubernetes, Prometheus, Thanos, VictoriaMetrics, Temporal, Cadence, Go, Python, Postgres, LLM APIs, MCP servers.
Responsibilities
Own the observability platform data pipelines and healthcheck frameworks; Define and build the unified API surface for infrastructure management; Build the production control plane for machine management and command execution; Maintain fleet state as the source of truth across provisioning and operations; Integrate new hardware generations (ZTP, DHCP, DNS) into the platform.
Seniority
Senior, hands-on IC