#1 Site Reliability Engineer II
Core
24/7 on-call Site Reliability Engineer responsible for real-time monitoring, incident triage, and root cause analysis in a FedRAMP Moderate cloud environment.
Role type
Mid-level SRE (IC)
Builds
Production cloud infrastructure and observability pipelines for government clients
Domain
Cloud Infrastructure / Cybersecurity Compliance (FedRAMP)
Deliverable
Production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux administration, Bash/Python scripting, Kubernetes/EKS, AWS core services, Log analysis (Kibana/Elasticsearch), CI/CD (GitLab/ArgoCD), Load balancing troubleshooting, Prometheus, Grafana
Preferred skills
CVE remediation, STIG hardening, ChatOps
Technologies
AWS, EKS, Kibana, Elasticsearch, GitLab CI/CD, ArgoCD, Argo Workflows, Prometheus, Grafana, Slack
Responsibilities
Real-time monitoring and first-response triage, Log-level root cause investigation, CVE remediation and STIG hardening, Deployment execution and troubleshooting
Seniority
Mid-level, hands-on IC