CareerPlanGet AI match score →

Site Reliability Engineer

San Francisco💼 Full-time🗓 2026-06-04 → 2026-07-31

Core

Lead scaling operational resilience, stability, observability, and debugging workflows for an AI workforce orchestration platform serving global enterprises.

Role type

Senior Site Reliability Engineer

Builds

Internal tooling for on-call teams, reliability workflows, and observability pipelines

Domain

AI infrastructure, distributed systems, enterprise operations

Deliverable

production ML models | infrastructure

Required skills

Go, Kubernetes, debugging production systems, observability tools (Grafana, Prometheus, Sentry), problem-solving under pressure

Preferred skills

distributed systems at scale, internal tooling development, CI/CD pipelines, infra-as-code, custom metrics and traces

Technologies

Go, Kubernetes, Grafana, Prometheus, Sentry

Responsibilities

Own system stability and uptime, design tools to reduce incident load, debug complex real-time failures, shift operations from reactive to proactive

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗