Site Reliability Engineer
Core
Define and monitor SLOs/SLIs, respond to production incidents, and maintain on-premises and cloud-based infrastructure to ensure system reliability and availability.
Role type
Site Reliability Engineer (SRE)
Builds
Production systems and automated software delivery pipelines
Domain
Cloud infrastructure and DevSecOps
Deliverable
production ML models | product features | infrastructure
Required skills
SLO/SLI definition, incident response, CI/CD pipeline design, infrastructure as code, container orchestration, scripting, cloud platform management, observability, security vulnerability assessment
Preferred skills
Capacity planning, system design for scalability, industry trend analysis
Technologies
Jenkins, GitLab CI, Harbor, Fortify, SonarQube, Blackduck, Docker, Kubernetes, Rancher, Bash, Python, Terraform, Ansible, AWS, Azure
Responsibilities
Define and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs); Respond to production incidents and conduct post-mortem analysis; Monitor system performance and troubleshoot issues; Design and maintain CI/CD pipelines; Implement and manage infrastructure as code (IaC); Implement security best practices and participate in vulnerability assessments; Manage and maintain on-premises and cloud-based infrastructure; Participate in capacity planning and system design
Seniority
Mid-level, hands-on IC