Site Reliability Engineer
Core
Establish SRE practices, frameworks, and feedback loops to make engineering teams more reliable across a global Postgres development platform.
Role type
Senior SRE Practice Lead / Reliability Expert
Builds
Operational readiness reviews, error budget policies, incident-to-improvement pipelines, and automation to reduce toil.
Domain
Cloud infrastructure, SRE discipline, Postgres platform operations
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
SRE practices design, SLO/SLI definition at scale, incident response facilitation, postmortem analysis, infrastructure-as-code, cloud infrastructure, automation development, multi-tenant system experience
Preferred skills
Kubernetes platform operations, OpenTelemetry, VictoriaMetrics, Grafana, developer-facing reliability tooling
Technologies
AWS, Pulumi, Terraform, CDK, Kubernetes, OpenTelemetry, VictoriaMetrics, Grafana
Responsibilities
Partner with service teams to define SLIs/SLOs and error budget policies; Own and evolve the Operational Readiness Review (ORR) process; Strengthen the incident-to-improvement pipeline; Act as reliability expert for architecture reviews and failure mode analysis; Identify and quantify operational toil and advocate for automation; Help teams design sustainable on-call practices; Track and report on org-wide operational maturity
Seniority
Senior, hands-on IC with strategic influence