CareerPlanGet AI match score →

Senior Site Reliability Engineer, AIOPs

US, CA, Santa Clara💼 Full-time💰 $148,000–$148,000🗓 2026-05-12 → 2026-07-31

Core

Operate an AI Data Center AIOps platform that transforms high-volume telemetry into reliable insights and automation for GPU fleets.

Role type

Senior Site Reliability Engineer (AIOps Platform)

Builds

Telemetry ingestion, processing, storage, APIs, and dashboards for GPU fleet management

Domain

AI Data Centers, GPU Infrastructure, Observability

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Kubernetes operations, distributed systems, incident response, SLO/SLI management, infrastructure as code, Python scripting, CI/CD pipelines, observability stack expertise

Preferred skills

Linux networking fundamentals, streaming systems operations (Kafka/Pulsar, Flink/Spark), automation tool development, large-scale cluster management

Technologies

Kubernetes, Terraform, Helm, Python, Bash, Prometheus, Grafana, Kafka, Pulsar, Flink, Spark, ClickHouse, Elastic, TSDBs

Responsibilities

Monitor platform health via dashboards/logs/metrics and automate recurring checks, own Kubernetes deployments end-to-end including runbooks and rollbacks, lead first-level incident triage and root cause analysis, build and maintain runbooks/SOPs/checklists, manage deployment infrastructure and packaging

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗