Senior Site Reliability Engineer
Core
Keep the Akuity GitOps platform running for enterprise customers, ensuring reliability across multi-region AWS infrastructure and shaping platform defense strategies.
Role type
Senior Site Reliability Engineer (SRE)
Builds
The Akuity SaaS platform (GitOps tools for Kubernetes)
Domain
Cloud Native / Kubernetes / SaaS
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (scheduler, networking, storage, autoscaling), AWS (EC2, EKS, VPC, S3, IAM), SLO/SLA definition and error budget management, observability (Prometheus, Grafana, OpenTelemetry), scripting/automation (Go, Python, Bash), incident command and post-mortem leadership
Preferred skills
Argo CD, Kargo, GitOps workflows, multi-region/multi-cluster Kubernetes, compliance infrastructure (SOC 2, ISO 27001, HIPAA, PCI DSS), developer tooling background
Technologies
Kubernetes (EKS), AWS, Argo CD, Kargo, Prometheus, Grafana, OpenTelemetry, Terraform
Responsibilities
Own SLI/SLO/SLA definitions and drive continuous improvement; Design, instrument, and maintain observability systems; Lead blameless post-mortems and close reliability gaps; Partner with engineering to build reliability into new features; Act as incident commander for high-severity events; Build and maintain runbooks and escalation paths
Seniority
Senior, hands-on IC