Senior Director, Cloud and Site Reliability Engineering
Core
Define and drive cloud infrastructure strategy and operational excellence for a SaaS platform, leading a team of Cloud Engineers and SREs to ensure high availability, reliability, and performance.
Role type
Senior Director, Cloud and Site Reliability Engineering
Builds
Cloud infrastructure roadmap, multi-cloud/hybrid-cloud architectures, SRE function (SLOs/SLIs/error budgets), and automated self-healing systems for a SaaS platform.
Domain
SaaS / Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure leadership, SRE principles (SLO/SLI/error budgets), incident management, infrastructure-as-code, Kubernetes, CI/CD, compliance frameworks, cost optimization, automation strategy
Preferred skills
AI and Agentic capabilities, multi-cloud strategy, regulated environment experience
Technologies
AWS, Azure, GCP, Kubernetes, Terraform, Pulumi
Responsibilities
Define and execute cloud infrastructure roadmap; Establish cloud architecture standards and best practices; Build and mature the SRE function; Own incident management and on-call strategy; Drive automation across infrastructure provisioning and observability; Partner with Security to ensure cloud environments meet compliance.
Seniority
Senior, hands-on IC with leadership responsibilities