Site Reliability Team Leader
Core
Lead a global Site Reliability Engineering (SRE) team to ensure the reliability, scalability, and operational excellence of a cloud-native marketing platform.
Role type
Senior IC SRE Team Lead (People Management + Technical Leadership)
Builds
Global SRE organization, automation tools, CI/CD pipelines, and observability infrastructure for a SaaS platform.
Domain
Cloud Infrastructure, SaaS, Marketing Technology
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes production operations, Cloud platform management (GCP/AWS), Python/Go/Bash scripting, CI/CD pipeline design, Observability stack implementation, Incident management, Distributed systems architecture, Linux networking, Remote team leadership
Preferred skills
Infrastructure as Code (Terraform/Ansible), Messaging systems (Kafka/Redis), Modern deployment strategies (Canary/Blue-Green), OpenTelemetry, SRE principles (SLIs/SLOs/Error Budgets), Large-scale SaaS experience, Cloud/Kubernetes certifications
Technologies
Kubernetes, GCP, AWS, Python, Go, Bash, Datadog, Prometheus, Grafana, Terraform, Ansible, Kafka, Redis, OpenTelemetry
Responsibilities
Manage and mentor a global SRE team across time zones, define technical roadmap for reliability and automation, oversee production rollouts and incident response, drive observability initiatives, and participate in on-call rotations.
Seniority
Senior, hands-on IC with people management responsibilities