Site Reliability Engineer
Core
Establish and maintain Site Reliability Engineering (SRE) practices, manage cloud infrastructure, and ensure system reliability and scalability.
Role type
Site Reliability Engineer
Builds
Cloud-based enterprise systems and infrastructure
Domain
Cloud Infrastructure (AWS/Azure)
Deliverable
infrastructure
Required skills
AWS cloud infrastructure, Azure, CI/CD, Infrastructure as Code (Terraform, Ansible, Packer, GitLab, Artifactory), Python, PowerShell, Bash, Linux, Windows, networking, incident management, root cause analysis, system design reviews, code reviews, Agile/Scrum/SAFe
Preferred skills
Government contract experience, RedHat/CompTIA certifications
Technologies
AWS, Azure, Terraform, Ansible, GitLab, Artifactory, Packer, Python, PowerShell, Bash
Responsibilities
Collaborate with cross-functional partners to establish SRE practices in Agile frameworks, participate in system design reviews to identify failure points and promote automation, conduct code reviews for efficiency and scalability, manage incident ceremonies to analyze root causes and mitigate downtime, create and maintain system documentation
Seniority
Mid-level, hands-on IC