#II Site Reliability Engineer II
Core
Real-time monitoring, first-response triage, log-level root cause investigation, CVE remediation, STIG hardening, and deployment execution on FedRAMP-authorised cloud platforms.
Role type
Site Reliability Engineer II (24/7 rotational shifts)
Builds
Production monitoring pipelines, security-hardened infrastructure, and CI/CD deployment workflows for government cloud environments
Domain
US Government Cloud (FedRAMP Moderate), AWS GovCloud
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux administration, Bash/Python scripting, Kubernetes/EKS management, AWS core services, L4/L7 load balancing troubleshooting, Prometheus, Grafana, GitLab CI/CD, ArgoCD, Argo Workflows, Kibana, Elasticsearch
Preferred skills
None stated
Technologies
AWS, EKS, GitLab CI/CD, ArgoCD, Argo Workflows, Kibana, Elasticsearch, Prometheus, Grafana, Slack
Responsibilities
Real-time monitoring and first-response triage, Log-level root cause investigation, CVE remediation and STIG hardening, Deployment execution and troubleshooting
Seniority
Mid-level (3+ years experience)