Staff Site Reliability Engineer, Federal (TS/SCI)
Core
Design, build, and operate large-scale, secure cloud infrastructure and production services for federal customers, ensuring high availability and reliability.
Role type
Staff Site Reliability Engineer (Technical Lead)
Builds
Highly reliable, scalable, and secure cloud services and automation platforms
Domain
Federal Government / Cloud Infrastructure / Security
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes production operations, Infrastructure as Code (Terraform, Helm), Go/Python, AWS/GCP, CI/CD, observability, incident response, SLO/SLI management, distributed data platforms
Preferred skills
SaaS platform operations, GitOps (ArgoCD), AI-assisted engineering tooling, Kubernetes microservices
Technologies
Go, Python, Terraform, Helm, Kubernetes, AWS, GCP, PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, ArgoCD
Responsibilities
Design and operate large-scale cloud infrastructure; lead incident response and post-incident reviews; define and improve SLIs/SLOs/error budgets; develop automation and self-service platforms; mentor engineers and guide operational best practices; drive projects from conception to production rollout.
Seniority
Staff, technical leadership & mentorship