SRE Engineer II
Core
Own L2 incident response, root cause analysis, and automation to ensure reliability of production infrastructure.
Role type
Senior Site Reliability Engineer (L2)
Builds
Production infrastructure reliability, automated operational tooling, and incident resolution workflows.
Domain
Cloud infrastructure (AWS/Azure/GCP) and network operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux OS administration, Cloud platform expertise (AWS/Azure/GCP), Python and Bash scripting, Network troubleshooting (TCP/IP, DNS, TLS), HTTP/HTTPS debugging, Observability stack usage (Prometheus, Grafana, ELK)
Preferred skills
Kubernetes, Docker, Message queues (Kafka, RabbitMQ), Database troubleshooting (MySQL, PostgreSQL, Redis), ITIL fundamentals
Technologies
Linux (RHEL/Ubuntu), AWS, Azure, GCP, Terraform, Python, Bash, Ansible, tcpdump, Wireshark, Nginx, HAProxy, Prometheus, Grafana, Datadog, ELK, Jira, ServiceNow
Responsibilities
Own L2 incident response and post-mortems; Monitor system health and maintain SLA/SLO adherence; Automate operational tasks to reduce toil; Collaborate on deployment reliability and capacity planning; Participate in on-call rotation and maintain runbooks
Seniority
Mid-Senior, hands-on IC