Senior Site Reliability Engineer
Core
Senior technical leader driving SRE principles, infrastructure reliability, and operational excellence for a global data resilience platform.
Role type
Senior Site Reliability Engineer (hands-on IC with mentorship)
Builds
Highly available, fault-tolerant, and scalable infrastructure on public clouds (Azure)
Domain
Cloud Infrastructure / Data Resilience / SaaS
Deliverable
production ML models | infrastructure
Required skills
Public cloud architecture (Azure), Distributed systems design, Infrastructure as Code (Terraform/Pulumi), Container orchestration (Kubernetes), Observability stack (Prometheus, Grafana, OpenTelemetry), CI/CD automation, Incident response and postmortems, Chaos engineering, Programming (JS, Node, TypeScript, Go, Java, C#)
Preferred skills
Large-scale B2B SaaS experience, Compliance frameworks (ISO, SOC 2, GDPR, FEDRAMP/CMMC)
Technologies
Azure, Kubernetes, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry, Node.js, TypeScript, Go, Java, C#
Responsibilities
Design and evolve cloud infrastructure for high availability and scalability; Establish and maintain SLIs, SLOs, and error budgets; Lead incident response and blameless postmortems; Drive adoption of deep observability practices; Develop automation and self-healing tools; Contribute to IaC, CI/CD, and deployment automation; Mentor engineers and advocate for DevOps/SRE best practices
Seniority
Senior, hands-on IC with mentorship