Senior Site Reliability Engineer
Core
Build and scale critical global infrastructure across data centers and cloud platforms to ensure stability and performance for every product.
Role type
Senior Site Reliability Engineer
Builds
Global compute platform, GitOps delivery pipelines, self-healing infrastructure, and internal tooling
Domain
Cloud infrastructure, distributed systems, gaming technology
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes cluster management, AWS cloud operations, Go and Python software development, Infrastructure as Code (IaC), Linux system administration, networking and packet-level debugging, container orchestration (Docker, containerd), capacity planning and autoscaling (Karpenter, HPA, KEDA), SLO definition and monitoring (Datadog)
Preferred skills
GCP, vSphere, Nutanix
Responsibilities
Drive stability and scalability across global compute platforms; Operate and evolve GitOps delivery models; Build self-healing infrastructure and internal tooling; Own cluster autoscaling and capacity strategy; Define SLOs and reliability metrics; Support technical growth through knowledge sharing and design discussions
Seniority
Senior, hands-on IC