Site Reliability Engineer
Core
Ensure world-class resilience and performance across the platform by advising on availability, scalability, observability, and capacity planning.
Role type
Senior Site Reliability Engineer
Builds
Highly available and resilient platform services
Domain
Cloud infrastructure and platform engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Performance monitoring and analysis, Capacity planning, Scripting and automation, Infrastructure as Code (Terraform), Relational database technologies (AWS Aurora), Messaging and distributed asynchronous workloads, Nginx, SRE processes, DevOps principles
Preferred skills
Enterprise solutions at scale, Containerization (Docker), Agile/Kanban, Software best practices (Refactoring, Clean Code, TDD)
Technologies
Terraform, DataDog, Prometheus, AWS Aurora, Docker, Nginx
Responsibilities
Proactively monitor and analyze platform performance, Collaborate with engineering teams to address performance bottlenecks, Implement and review SLOs, Improve observability through monitoring, alerting, and dashboards, Ensure service high availability and resilience, Devise runbooks and test DR plans, Conduct capacity assessments and scaling plans, Respond to and troubleshoot incidents, Participate in blameless postmortems, Develop and maintain playbooks and documentation
Seniority
Senior, hands-on IC