Cleared Senior Site Reliability Engineer
Core
Build and operate production systems for AI-driven logistics decision platforms in classified, air-gapped, and disconnected environments.
Role type
Senior Site Reliability Engineer (Infrastructure & Automation)
Builds
Production systems, CI/CD pipelines, monitoring/observability pipelines, and automated operational workflows for logistics decision platforms.
Domain
National Security / Defense Logistics / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, networking, cloud infrastructure (AWS/Azure/GovCloud), infrastructure-as-code (Terraform/Ansible), container orchestration (Kubernetes/Docker), CI/CD pipeline design, incident response, observability tooling (Prometheus/Grafana/ELK), SLO/SLI definition, documentation.
Preferred skills
Azure Government (GCC High/Secret) experience, Authority to Operate (ATO) under RMF, multi-cloud operations, edge/DDIL deployment patterns, startup environment experience, logistics/supply chain domain knowledge.
Responsibilities
Own reliability, availability, and performance of production systems in classified environments; build and maintain monitoring/alerting pipelines; lead incident response and postmortems; design and operate CI/CD pipelines and IaC; automate manual operational work; partner with engineers to define SLOs/SLIs; work with government customers to translate operational constraints into system requirements.
Seniority
Senior, hands-on IC