Site Reliability Engineer
Core
Ensure the availability, reliability, and performance of production systems through monitoring, incident management, and system health checks.
Role type
Site Reliability Engineer (SRE)
Builds
Production systems, databases, APIs, and cloud infrastructure
Domain
Cloud infrastructure and systems engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux/Unix fundamentals, AWS or GCP, Grafana, Prometheus, CloudWatch, New Relic, ELK, Loki, Wazuh, MySQL, Aurora DB, MongoDB, DNS, TCP/IP, Load Balancers, Jira, PagerDuty
Preferred skills
None stated
Technologies
AWS, GCP, Grafana, Prometheus, CloudWatch, New Relic, ELK, Loki, Wazuh, MySQL, Aurora DB, MongoDB, Jira, PagerDuty
Responsibilities
Monitor servers, applications, databases, cloud infrastructure, and third-party integrations; Configure and manage alerts, dashboards, and observability tools; Respond to incidents, perform initial diagnosis, and coordinate escalations; Analyze system and application logs to identify issues; Perform daily health checks for production systems; Maintain SOPs, runbooks, incident reports, and shift handover reports; Provide L1/L2 production support and assist engineering teams with troubleshooting
Seniority
Mid-level (1-4 years experience)