Site Reliability Engineer (SRE) – II
Core
Maintain 24/7 platform availability, hardware reliability, and security posture for a hybrid environment spanning bare-metal RHEL servers, AWS (Commercial and GovCloud), and Kubernetes (EKS) under FedRAMP standards.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Hybrid cloud and on-premises infrastructure ensuring continuous uptime and compliance (via careerplan.io/jobs/RP1038871-site-reliability-engineer-sre-ii-at-ffive)
Domain
Cybersecurity, Cloud Infrastructure, Hybrid Operations
Required skills
Bare-metal Linux administration (RHEL), Kubernetes (EKS) operations, AWS (Commercial/GovCloud) management, Prometheus/Grafana monitoring, Elasticsearch/Kibana log analysis, GitOps (ArgoCD/GitLab CI/CD), Terraform, Network load balancing (L4/L7), Incident response
Preferred skills
Shell scripting (Bash/Python), DISA STIG hardening, FIPS 140 compliance, Vendor hardware coordination
Technologies
RHEL, AWS, EKS, Kubernetes, Prometheus, Grafana, Elasticsearch, Kibana, GitLab CI/CD, ArgoCD, Argo Workflows, Terraform, Helm, Kustomize
Responsibilities
Perform 24/7 eyes-on-glass monitoring and incident triage; Troubleshoot bare-metal server hardware and OS issues; Execute FedRAMP security controls and vulnerability patching; Manage EKS clusters and AWS networking components; Operate CI/CD pipelines and GitOps deployments; Document incidents and automate operational tasks
Seniority
Mid-Senior, hands-on IC
