Staff Site Reliability Engineer - Site Experience
Core
Lead reliability engineering for critical user-facing systems at internet scale, ensuring fast, reliable, and resilient interactions across web, mobile, APIs, feeds, and real-time systems.
Role type
Staff Site Reliability Engineer (Infrastructure & Product Intersection)
Builds
High-availability, performant distributed systems for Reddit's global user experience
Domain
Internet-scale social media platform infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems architecture, Linux systems administration, Cloud native architectures, Go/Python programming, Observability (metrics/logging/tracing), SLO/SLI definition, Incident management, Capacity planning, Automation engineering
Preferred skills
Kubernetes, CDN optimization, Traffic engineering, Edge reliability, Open source contributions
Technologies
Go, Python, Kubernetes, Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis
Responsibilities
Drive reliability and scalability for critical user-facing services; Architect highly available systems under massive global load; Reduce operational risk through proactive mitigation; Build automation tooling for deployment safety and incident response; Lead complex incident response and postmortems; Mentor engineers and define reliability engineering standards
Seniority
Staff, technical leadership & mentorship