Staff Site Reliability Engineer - Site Experience
Core
Lead reliability engineering initiatives for critical user-facing systems at internet scale, ensuring fast, reliable, and resilient interactions across web, mobile, APIs, feeds, and real-time systems.
Role type
Staff Site Reliability Engineer (IC with leadership)
Builds
High-availability, performant systems for Reddit's global user base
Domain
Internet-scale consumer social platform
Deliverable
production ML models | product features | infrastructure
Required skills
distributed systems, networking, Linux systems, cloud native architectures, Go/Python programming, observability (metrics/logging/tracing), SLOs, automation, incident management, capacity planning, architectural design for scale
Preferred skills
Kubernetes, containers, CDN optimization, edge reliability, traffic engineering, global infrastructure, open source contributions
Technologies
Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis
Responsibilities
Drive reliability and scalability for critical user-facing services; Architect highly available systems under massive global load; Identify and mitigate systemic risks; Build automation for deployment safety and incident response; Lead complex incident response and postmortems; Mentor engineers and define reliability engineering best practices
Seniority
Staff, hands-on IC with strategic influence