Principal Site Reliability Engineer
Core
Define and execute long-term strategy for Kubernetes platform across GKE, EKS, RKE2, and on-premise environments to ensure reliability and scalability for sports betting and gaming platforms.
Role type
Principal Site Reliability Engineer (Infrastructure Strategy)
Builds
Cloud and on-premise Kubernetes clusters, automation-first infrastructure, and self-healing systems for high-scale gaming platforms.
Domain
Sports betting and gaming industry; Cloud infrastructure and distributed systems.
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes architecture, Infrastructure as Code (Terraform/Pulumi), Go/Python, GitOps, observability, distributed systems, Linux, capacity planning, cost optimization, incident management, technical leadership, mentoring.
Preferred skills
Regulated industry experience, hybrid cloud environments, open-source contributions, cloud certifications.
Technologies
Google Kubernetes Engine, Amazon Elastic Kubernetes Service, RKE2, Terraform, Pulumi, Go, Python, Linux.
Responsibilities
Drive architectural decisions for critical infrastructure; Lead large-scale platform initiatives; Establish reliability practices (SLOs, error budgets); Build automation tooling; Champion AI-powered engineering capabilities; Lead critical platform incidents; Mentor senior engineers.
Seniority
Principal, hands-on IC with strategic leadership