Site Reliability Engineer
Core
Design, deploy, and support cloud-based systems, focusing on automation, incident management, and system reliability within enterprise environments.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud-based systems, automated infrastructure, and self-healing operational processes
Domain
Enterprise IT, Cloud Infrastructure (AWS/Azure)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS cloud infrastructure, Azure, CI/CD pipelines, Infrastructure as Code (Terraform, Ansible), Python scripting, Linux/Windows administration, networking fundamentals, incident management, root cause analysis, Agile methodologies (Scrum, Kanban), source control best practices
Preferred skills
Public Trust clearance, DoD experience
Technologies
AWS, Ansible, Azure, Bash, GitLab, Linux, PowerShell, Python, Terraform, Windows
Responsibilities
Establish and sustain SRE practices in Agile environments, participate in system design reviews to identify failure points, conduct code reviews for efficiency and scalability, manage incidents to determine root causes and reduce downtime, create and maintain operational documentation
Seniority
Mid-level, hands-on IC