Site Reliability Engineer
Core
Own application-level infrastructure and reliability for a high-traffic commerce domain, ensuring the backbone that lets game developers get paid.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Deploy pipelines, Kubernetes manifests, SLOs, capacity planning, and production readiness for commerce services.
Domain
Video game industry commerce and monetization
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (Helm, manifests), observability (SLOs/SLIs, Datadog, OpenTelemetry), software engineering (Go, PHP, Python, Bash), Infrastructure as Code (Terraform), CI/CD (GitLab CI, GitHub Actions), incident response, capacity planning, automation scripting
Preferred skills
Kubernetes certifications, GCP certifications, HashiCorp certifications, payments/fintech/e-commerce/gaming domain experience
Technologies
Kubernetes, GCP/GKE, Datadog, OpenTelemetry, Terraform, GitLab CI, GitHub Actions, Helm
Responsibilities
Own application-level infrastructure including Helm charts, Terraform, and Kubernetes deployments; design and implement observability (SLOs/SLIs, monitors, alerts); evolve CI/CD pipelines; perform capacity planning and performance tuning; run Production Readiness Reviews; support incident response and post-mortems; build domain-specific automation; drive reliability roadmap; participate in product planning and architecture reviews; co-author company-wide SRE standards; participate in SRE duty rotation
Seniority
Senior, hands-on IC