Senior Site Reliability Engineer
Core
Operate, maintain, upgrade, and improve production Kubernetes clusters and workloads across AWS and Azure, while supporting distributed data systems and building automation tooling.
Role type
Senior Site Reliability Engineer (IC)
Builds
Production Kubernetes clusters, distributed data systems, and automation tooling for AWS and Azure environments.
Domain
Cloud Infrastructure (AWS, Azure) and Container Orchestration (Kubernetes)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes cluster management, AWS and Azure cloud infrastructure, Terraform, Python/Go/Bash scripting, Linux systems administration, incident response, capacity planning, observability, networking, storage, and high-availability patterns.
Preferred skills
Experience with Helm, Kustomize, operators, Istio, autoscaling, and disaster recovery initiatives.
Technologies
Kubernetes, Helm, Kustomize, Istio, Terraform, Python, Go, Bash, AWS, Azure, ClickHouse, Aurora
Responsibilities
Operate and maintain production Kubernetes clusters; manage cluster and node lifecycle; support production database platforms; build and maintain infrastructure using Terraform; develop automation and tooling; participate in on-call rotations and lead incident resolution; improve observability and monitoring; partner with engineering teams on scalable service design; contribute to capacity planning and resilience initiatives.
Seniority
Senior, hands-on IC
