Senior Site Reliability Engineer
Core
Build, operate, and improve production systems and internal platforms supporting SolarWinds products.
Role type
Senior Site Reliability Engineer (IC)
Builds
Production Kubernetes clusters, workloads, and database platforms on AWS and Azure
Domain
Cloud Infrastructure / Site Reliability Engineering
Required skills
Kubernetes, AWS, Azure, Linux, Terraform, Python, Go, Bash, Incident Management, Observability, Capacity Planning
Responsibilities
Operate and upgrade production Kubernetes clusters and workloads across AWS and Azure; Manage Kubernetes platform components including Helm, Kustomize, operators, and Istio; Support and maintain production database platforms like ClickHouse and Aurora; Build and maintain infrastructure using Terraform; Develop automation and tooling using Python, Go, or Bash; Participate in on-call rotation and lead incident resolution; Improve observability, monitoring, and logging; Partner with engineering teams to design scalable services; Contribute to capacity planning and infrastructure lifecycle management.
