Site Reliability Engineer
Core
Design, build, and maintain scalable, high-performing cloud infrastructure across AWS GovCloud and Azure platforms while establishing high availability and automating incident response.
Role type
Site Reliability Engineer (Cloud Infrastructure)
Builds
Mission-critical SaaS and PaaS platforms for construction, geospatial, and transportation industries
Domain
Cloud Infrastructure, AWS GovCloud, FedRAMP Compliance
Deliverable
production ML models | product features | infrastructure
Required skills
AWS and/or Azure public cloud administration, Python/PowerShell/Bash/Perl scripting, IT operations and networking concepts, automated incident response, SLO definition, redundancy and failover testing
Preferred skills
Infrastructure-as-code (Ansible, Terraform), container platforms (Docker, ECS, EKS, Kubernetes), FedRAMP frameworks, AWS SDK, directory services (SAML, Active Directory), database technologies (RDBMS/NoSQL), monitoring tools (Grafana, Sumo, PRTG)
Technologies
AWS GovCloud, Azure, Python, PowerShell, Bash, Perl, Ansible, Terraform, Docker, ECS, EKS, Kubernetes, Grafana, Sumo, PRTG, Jira Service Management
Responsibilities
Design and maintain scalable cloud infrastructure; Architect automated deployment, monitoring, and observability tools; Collaborate on SLOs and automated incident response systems; Conduct redundancy, resilience, and failover testing; Manage security posture and enforce FedRAMP best practices
Seniority
Mid-level, hands-on IC
