Site Reliability Engineering Manager
Core
Lead reliability, scalability, performance, and operational excellence for critical production platforms, combining deep database expertise with broader service ownership.
Role type
Senior Site Reliability Engineering Manager (Database & Platform)
Builds
Highly available, measurable, automated, and resilient database and platform services
Domain
Cloud infrastructure, relational databases, and production engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Site Reliability Engineering, Database Engineering, Team Leadership, PostgreSQL, Oracle, SQL Server, Cloud Operations, Automation, Observability, Incident Response, Performance Tuning, High Availability, Disaster Recovery, Kubernetes, SLO/SLI Definition, Cost Optimization
Preferred skills
Infrastructure as Code, Self-healing mechanisms, Capacity Planning, Architecture Reviews, Release Management
Technologies
PostgreSQL, Oracle, SQL Server, AWS, Kubernetes
Responsibilities
Lead and develop a team of reliability engineers; Define reliability strategy, operating model, and roadmap; Own availability, performance, scalability, and resilience of production databases; Define and enforce SLIs, SLOs, and error budgets; Lead incident management and postmortems; Drive Infrastructure as Code and operational automation; Partner with Development, Platform Engineering, Infrastructure, Security, and Product teams.
Seniority
Senior, hands-on IC with leadership responsibilities
