CareerPlanSign in

Distributed Systems Engineer

San Francisco, CA💼 Full-time🗓 2026-07-05 → 2026-09-26

Core

Building and operating the observability platform, control plane, and fleet state management for a hyperscale AI compute infrastructure.

Role type

Senior IC distributed systems engineer (infrastructure)

Builds

Production observability platform, unified API surface for infrastructure, Kubernetes-based control plane, and fleet state source of truth.

Domain

AI infrastructure, distributed systems, data centers

Deliverable

production ML models | infrastructure

Required skills

Distributed systems, API design and versioning, Kubernetes, observability stacks (Prometheus, Thanos, VictoriaMetrics), workflow orchestration (Temporal, Cadence), hardware telemetry (BMC/Redfish), Go, Python, Postgres

Preferred skills

Time-series data engineering, ZTP/DHCP/DNS implementation, AI tooling (LLM APIs, MCP servers, agentic frameworks)

Responsibilities

Build and operate data pipelines and correlation engines for fleet observability; Design and maintain the unified API surface for infrastructure management; Implement the production control plane for machine management and command execution; Ensure fleet state matches reality across provisioning and operations; Integrate new hardware generations and sites into the platform.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.