Senior Site Reliability Engineer
Core
Design, build, and operate scalable multi-cloud and hybrid infrastructure for 2K's global game services, ensuring millions of players stay connected during live events.
Role type
Senior Site Reliability Engineer (Infrastructure & Platform)
Builds
Production game services, account platforms, CI/CD pipelines, and developer tooling across AWS, GCP, and on-premises data centers.
Domain
Video Game Industry / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes (EKS/GKE), Infrastructure as Code (Terraform/Pulumi), Observability (Prometheus/Grafana/Datadog), Linux internals, Networking (TCP/IP/DNS/TLS), Incident Management, Automation, Go/Python/TypeScript
Preferred skills
Service mesh (Istio/Cilium), FinOps, AI/Agentic Development, Mentoring
Technologies
Terraform, Pulumi, ArgoCD, Flux, EKS, GKE, Istio, Cilium, Prometheus, Grafana, Datadog, GitHub Actions, Jenkins, OPA, Gatekeeper, Ansible, Puppet, AWS Systems Manager
Responsibilities
Design and operate multi-cloud/hybrid infrastructure; Own Kubernetes cluster lifecycle and networking; Push progressive delivery patterns; Build and run observability stack; Lead chaos engineering and incident response; Eliminate toil via automation; Promote SRE practices and shape architectural decisions.
Seniority
Senior, hands-on technical leader
