CareerPlanSign in

Staff Site Reliability Engineer

💼 Full-time🗓 2026-09-18 → 2026-09-26

Core

Architect and implement observability, drive automation, lead incident response, and optimize Kubernetes-based infrastructure to ensure high availability and scalability for Replit's platform serving millions of developers.

Role type

Staff Site Reliability Engineer (IC)

Builds

Production ML models | product features | dashboards & analysis | infrastructure

Domain

Cloud-native infrastructure, distributed systems, developer tools

Deliverable

infrastructure

Required skills

Python, Go, Kubernetes, distributed systems, observability (metrics/logging/tracing), incident management, infrastructure as code (Terraform/Pulumi), capacity planning, debugging complex systems

Preferred skills

Google Cloud Platform (GCP), Prometheus, Grafana, Datadog, OpenTelemetry, rapid-growth startup experience, technical writing

Technologies

Kubernetes, Docker, GCP, Terraform, Pulumi, Python, Go, Prometheus, Grafana, Datadog, OpenTelemetry

Responsibilities

Design and lead implementation of comprehensive monitoring, logging, and tracing solutions; Define and drive reliability standards including SLOs and SLIs; Lead incident management and conduct blameless post-mortems; Architect and build automation to eliminate toil and improve CI/CD pipelines; Optimize performance on large-scale Kubernetes deployments; Debug and harden distributed systems across the stack; Provide staff-level guidance on system designs for reliability and scalability; Educate and mentor the broader engineering team on reliability practices

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.