Distributed Systems Engineer
Core
Building and operating the observability platform, control plane, and fleet state management for a hyperscale AI compute infrastructure.
Role type
Senior IC distributed systems engineer (infrastructure)
Builds
Production observability platform, unified API surface for infrastructure, Kubernetes-based control plane, and fleet state source of truth.
Domain
AI infrastructure, distributed systems, data centers
Deliverable
production ML models | infrastructure
Required skills
Distributed systems, API design and versioning, Kubernetes, observability stacks (Prometheus, Thanos, VictoriaMetrics), workflow orchestration (Temporal, Cadence), hardware telemetry (BMC/Redfish), Go, Python, Postgres
Preferred skills
Time-series data engineering, ZTP/DHCP/DNS implementation, AI tooling (LLM APIs, MCP servers, agentic frameworks)
Responsibilities
Build and operate data pipelines and correlation engines for fleet observability; Design and maintain the unified API surface for infrastructure management; Implement the production control plane for machine management and command execution; Ensure fleet state matches reality across provisioning and operations; Integrate new hardware generations and sites into the platform.
Seniority
Senior, hands-on IC