Senior Site Reliability Engineer (SRE) – Technical Leader, Kubernetes Platform (IRAP)
Core
Design, operate, and scale a Kubernetes-based platform supporting various environments to ensure reliability, observability, and compliance.
Role type
Senior IC SRE Technical Leader (Kubernetes Platform)
Builds
Production-grade Kubernetes platforms, CI/CD pipelines, and observability stacks
Domain
Cloud Infrastructure / Kubernetes / Platform Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (EKS/AKS/GKE), Linux systems, networking fundamentals, Infrastructure as Code (Terraform), CI/CD systems, Python/Go scripting, observability platforms (Prometheus/Grafana/OpenTelemetry/ELK), compliance frameworks (PCI/ISO)
Preferred skills
Technical direction influence, cross-functional initiative driving, engineering mentorship
Technologies
Kubernetes, Terraform, GitHub Actions, GitLab CI, Jenkins, ArgoCD, Prometheus, Grafana, OpenTelemetry, ELK
Responsibilities
Design and operate production-grade Kubernetes platforms, define and drive SLOs/SLIs/error budgets, build and evolve secure CI/CD pipelines, implement robust observability, reduce operational toil via automation, partner with security/compliance teams, support audit processes, participate in on-call rotations and incident response
Seniority
Senior, hands-on IC with leadership responsibilities