Site Reliability Egineer
Core
Maintain and improve observability systems, define Service Level Objectives, and drive proactive reliability improvements for high-traffic, public-facing platforms.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production platforms and infrastructure for high-traffic, public-facing services
Domain
Cloud infrastructure and DevOps
Deliverable
production ML models | infrastructure
Required skills
Incident management, scripting (Python, Bash, PowerShell), CI/CD tools, monitoring/APM tools, containerization, serverless services
Preferred skills
Infrastructure as Code, runbook creation, failure scenario analysis, e-commerce domain experience
Technologies
GitHub Actions, Azure DevOps, GitLab, Jenkins, Datadog, New Relic, Dynatrace, Prometheus, Grafana, AWS, Azure, GCP, Docker, Ansible, Chef
Responsibilities
Maintain and improve observability systems (monitoring, logging, alerting); Define, adjust, and maintain Service Level Objectives (SLOs); Participate in incident resolution and on-call rotations; Drive proactive reliability improvements across platforms; Collaborate with teams to analyze failure scenarios and implement mitigations; Create and maintain runbooks for incident response and prevention; Eliminate non-value-adding tasks through automation and process optimization
Seniority
Senior, hands-on IC