Forward Deployed Site Reliability Engineer (TS/SCI Required)
Core
Forward-deployed SRE ensuring reliability and performance of mission-critical platforms in restricted, air-gapped government environments.
Role type
Senior Site Reliability Engineer (Forward Deployed)
Builds
Mission-critical cyber operations platform for U.S. government and allied customers
Domain
Cybersecurity / Defense / Government Contracting
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE/production operations, SLIs/SLOs/error budgets, Docker/Docker Compose, AWS (EC2/ECS/RDS/VPC), Linux/Unix administration, Terraform, LGTM stack (Grafana/Loki/Tempo/Mimir), incident response, Python/Bash scripting
Preferred skills
Go, NATS pub/sub, cyber operations/intelligence background, AWS certifications
Technologies
AWS, Docker, Terraform, Grafana, Loki, Tempo, Mimir, Python, Bash, Go
Responsibilities
Define and track SLIs/SLOs and use error budgets to drive reliability; lead on-site incident response and triage; maintain observability dashboards and alerting; automate operational toil; manage containerized service deployments and rollbacks; perform capacity planning; serve as primary technical liaison between on-site operations and remote engineering teams.
Seniority
Senior, hands-on IC with customer-facing responsibilities