Senior SRE Engineer
Core
Own and mature reliability, performance, and security across a growing multi-cluster Kubernetes platform serving dozens of consumer AI apps globally.
Role type
Senior Site Reliability Engineer (Infrastructure & Security)
Builds
Multi-cluster Kubernetes (GKE) environment, observability stack, and security-hardened platform tooling
Domain
Consumer AI applications, Cloud Infrastructure, Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Kubernetes operations, SLO/SLI definition, observability, Infrastructure as Code, incident command, security hardening, automation scripting, cloud networking
Preferred skills
GitOps, chaos engineering, FinOps, service mesh, compliance frameworks
Technologies
Kubernetes, GCP, Terraform, Cloudflare, CI/CD pipelines
Responsibilities
Define and enforce SLIs/SLOs/error budgets, operate multi-cluster Kubernetes, lead incident response, build self-service platform tooling, embed security practices, reduce operational toil through automation
Seniority
Senior, hands-on IC
