CareerPlanGet AI match score →

Staff Site Reliability Engineer

💼 Full-time🗓 2026-05-20 → 2026-07-31

Core

Architect and maintain resilient, scalable infrastructure for a global agentic software creation platform serving millions of developers.

Role type

Staff Site Reliability Engineer (IC)

Builds

Production-grade observability solutions, automation frameworks, and self-healing systems for a cloud-native platform.

Domain

Cloud-native infrastructure, distributed systems, and developer tools.

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Python or Go, Kubernetes, distributed systems architecture, observability (metrics/logging/tracing), incident management, Infrastructure as Code (Terraform/Pulumi), capacity planning, debugging complex distributed systems.

Preferred skills

Google Cloud Platform (GCP), Prometheus, Grafana, Datadog, OpenTelemetry, rapid-growth startup environment experience.

Technologies

Kubernetes, Terraform, Pulumi, Python, Go, GCP, Prometheus, Grafana, Datadog, OpenTelemetry, Docker.

Responsibilities

Design and lead implementation of comprehensive monitoring, logging, and tracing solutions; define and drive Service Level Objectives (SLOs) and Service Level Indicators (SLIs); lead incident response and conduct blameless post-mortems; architect and build automation to eliminate toil; optimize performance on large-scale Kubernetes deployments; debug and harden distributed systems; review feature and system designs for reliability and scalability; mentor and educate the broader engineering team.

Seniority

Staff, hands-on IC with strategic guidance and mentorship responsibilities.

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗