Staff Site Reliability Engineer
Core
Design, build, and operate secure, resilient, scalable cloud infrastructure and developer tools for large-scale customer-facing products.
Role type
Senior IC Staff Site Reliability Engineer (Infrastructure)
Builds
Production infrastructure systems, developer tooling, and observability platforms
Domain
Cloud infrastructure, distributed systems, reliability engineering
Required skills
Cloud architecture, Terraform, Kubernetes/EKS, Go/Python, Redis/ElastiCache, Prometheus/Grafana/OpenTelemetry, incident response, performance tuning, security mindset, code reviews, technical documentation
Preferred skills
AWS experience, 6-10 years in infrastructure/platform/backend engineering, experience owning full lifecycle of infrastructure products, T-shaped technical expertise
Technologies
AWS, Terraform, Kubernetes, EKS, Go, Python, Redis, ElastiCache, Prometheus, Grafana, OpenTelemetry
Responsibilities
Take end-to-end ownership of core infrastructure subsystems from design to production operations; Define project goals and success metrics; Translate requirements into practical designs; Build secure, reliable, high-performing infrastructure; Develop and deploy production software and developer tools; Manage infrastructure via code and configuration; Partner with product engineering to design for scale; Participate in incident response and debugging; Develop monitoring and observability practices; Apply security-focused mindset; Mentor engineers and drive collaboration
Seniority
Senior, hands-on IC with mentorship responsibilities
