Director, Site Reliability Engineering
Core
Foundational leader to build and lead Site Reliability Engineering (SRE) functions from the ground up in Dublin, ensuring secure reliability, scalability, and performance for Klaviyo's global SaaS platform.
Role type
Director, Site Reliability Engineering (Foundational Leader)
Builds
Unified SRE organization, automated self-service tooling, and infrastructure enabling fast, reliable deployments for 200,000+ businesses.
Domain
SaaS, Cloud-Native Infrastructure, Distributed Systems, Data Processing Pipelines, AI Services
Deliverable
production ML models | infrastructure
Required skills
Cloud-native production systems at scale, Distributed systems, Observability, Container orchestration (Kubernetes), Infrastructure as code (Terraform), CI/CD principles, Geographically distributed team management, GDPR compliance, Incident command, Strategic roadmap definition
Preferred skills
AI experimentation and fluency, Building new offices/teams from scratch
Technologies
AWS, Kubernetes, Terraform
Responsibilities
Define vision and strategy for the SRE organization; Build and mentor a multi-disciplinary team in Dublin; Embed reliability principles into the software development lifecycle; Own platform availability, latency, performance, and capacity planning; Develop automated, self-service tooling; Ensure global compliance with GDPR and data protection regulations; Command reliability incidents and lead blameless post-mortems.
Seniority
Director, Hands-on IC with strategic leadership