Staff Site Reliability Engineer
Core
Architect and implement observability, drive automation, lead incident response, and optimize Kubernetes-based infrastructure to ensure high availability and scalability for Replit's platform serving millions of developers.
Role type
Staff Site Reliability Engineer (IC)
Builds
Production ML models | product features | dashboards & analysis | infrastructure
Domain
Cloud-native infrastructure, distributed systems, developer tools
Deliverable
infrastructure
Required skills
Python, Go, Kubernetes, distributed systems, observability (metrics/logging/tracing), incident management, infrastructure as code (Terraform/Pulumi), capacity planning, debugging complex systems
Preferred skills
Google Cloud Platform (GCP), Prometheus, Grafana, Datadog, OpenTelemetry, rapid-growth startup experience, technical writing
Technologies
Kubernetes, Docker, GCP, Terraform, Pulumi, Python, Go, Prometheus, Grafana, Datadog, OpenTelemetry
Responsibilities
Design and lead implementation of comprehensive monitoring, logging, and tracing solutions; Define and drive reliability standards including SLOs and SLIs; Lead incident management and conduct blameless post-mortems; Architect and build automation to eliminate toil and improve CI/CD pipelines; Optimize performance on large-scale Kubernetes deployments; Debug and harden distributed systems across the stack; Provide staff-level guidance on system designs for reliability and scalability; Educate and mentor the broader engineering team on reliability practices
Seniority
Senior, hands-on IC with mentorship responsibilities