Senior Manager, Site Reliability Engineering (Federal)
Core
Leading a team of SREs to scale high-throughput, 99.999 availability identity infrastructure on AWS, focusing on Edge networking, K8s, CI/CD, and observability.
Role type
Senior Manager, Site Reliability Engineering
Builds
Scalable, cost-effective, and efficient cloud infrastructure and tooling for the IDaaS platform
Domain
Identity & Access Management (IAM), Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Technical leadership, people management, Agile/DevOps methodologies, cloud-native architectures, containerization (Kubernetes), Infrastructure as Code (Terraform), CI/CD pipelines, observability platforms (Grafana, Splunk, APM), software development, PaaS, automation
Preferred skills
Multi-cloud environment experience
Technologies
AWS, Kubernetes, Terraform, Grafana, Splunk, APM
Responsibilities
Manage a team of SREs supporting workloads in private sector environments; Drive microservice journey, DevOps maturity, and workload reliability; Accelerate SRE and product engineering velocity via tooling and self-service capabilities; Lead, mentor, and grow high-performing engineering teams; Perform engineering design evaluations and ensure project completion within constraints; Improve SDLC processes for Cloud infrastructure as code; Manage service and business expectations and prioritize resource allocation
Seniority
Senior, hands-on IC with management responsibilities