Senior DevOps SRE Engineer
Core
Own and evolve AWS infrastructure to keep mission-critical 911 systems reliable, scalable, and secure for public safety agencies.
Role type
Senior Site Reliability Engineer (Infrastructure & Platform)
Builds
AWS environments, containerized workloads, CI/CD pipelines, and self-service internal developer platforms.
Domain
Public Safety / Emergency Response / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Kubernetes, Terraform, Bash scripting, SRE principles (SLOs/SLIs/error budgets), incident management, distributed systems, project leadership
Preferred skills
AI-assisted engineering tools (Claude, MCP), AI/ML infrastructure (model deployment, inference pipelines)
Technologies
AWS, Terraform, Terragrunt, Kubernetes, Docker, Argo, Datadog, Prometheus, Grafana, Bitbucket, Jenkins, JIRA
Responsibilities
Own and evolve AWS infrastructure using IaC; Architect and scale AWS environments; Deploy, scale, and manage containerized workloads; Lead deployment and release processes; Define and enforce SLOs, SLIs, and error budgets; Drive full utilization of Datadog for monitoring; Build self-service internal developer platforms; Partner cross-functionally on long-term technical planning; Document work and provide cross-training; Resolve JIRA tickets across Cloud, CI/CD, deployments, and monitoring.
Seniority
Senior, hands-on IC