Staff Site Reliability Engineer
Core
Evolving Site Reliability Engineering into a Platform as a Service (PaaS) organization to host AI-driven internal workflows reliably at scale.
Role type
Staff Site Reliability Engineer (Platform Engineering)
Builds
Internal developer platform (Kubernetes-based PaaS) for AI/ML workflows, model serving, and GPU scheduling
Domain
Cloud-native infrastructure, Kubernetes, AI/ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes architecture at production scale, AWS public cloud, Infrastructure as Code (Terraform), CI/CD pipelines, Observability (Grafana, Splunk, APM), Technical leadership across geographies, Mentoring engineers
Preferred skills
Internal Developer Platforms (IDP), AI agent orchestration, Vector databases, Service mesh (Istio, Linkerd), GitOps (ArgoCD, Flux), SCM platforms (GitHub, GitLab)
Technologies
Kubernetes, AWS, Terraform, Grafana, Splunk, Istio, ArgoCD, GitHub
Responsibilities
Owning architecture and evolution of the Kubernetes platform for Dublin/EMEA, Leading technical design of self-service platform capabilities, Driving readiness for AI-driven workflows, Taking ownership of complex production incidents and scaling challenges, Setting technical standards for SLIs/SLOs and on-call practices, Mentoring engineers and driving technical alignment
Seniority
Staff, hands-on IC with strategic leadership
