Senior Site Reliability Engineer
Core
Design, build, and operate scalable multi-cloud and hybrid infrastructure for global video game services, ensuring millions of players stay connected during live events.
Role type
Senior Site Reliability Engineer (Infrastructure Owner)
Builds
Multi-cloud game services, account platforms, CI/CD pipelines, and developer tooling across AWS, GCP, and on-premises data centers.
Domain
Video Games / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes (EKS/GKE), Infrastructure as Code (Terraform/Pulumi), Observability (Prometheus/Grafana/Datadog), Linux internals, TCP/IP networking, Incident response, Go/Python/TypeScript, Service Mesh (Istio/Cilium)
Preferred skills
Live-service game experience, FinOps, AI/Agentic Development, Cloud certifications, Mentoring
Technologies
AWS, GCP, VMware, Kubernetes, Terraform, Pulumi, ArgoCD, Flux, Istio, Cilium, Prometheus, Grafana, Datadog, OpenTelemetry, GitHub Actions, Jenkins, Ansible, Puppet, AWS Systems Manager, PasswordState, 1Password, AWS Secrets Manager, OPA, Gatekeeper
Responsibilities
Design and operate scalable multi-cloud infrastructure; Define SLI/SLO/error budget policies; Lead chaos engineering exercises; Drive incident response and post-mortems; Eliminate toil through automation; Promote SRE practices across studios.
Seniority
Senior, hands-on technical leader