SRE Engineer II/III
Core
Maintain reliability, availability, and performance of production environments; serve as escalation point for incidents; drive root cause analysis and reduce toil through automation.
Role type
L2/L3 Site Reliability Engineer
Builds
Production system stability, automated operational tooling, incident response workflows
Domain
Cloud infrastructure (AWS/Azure/GCP), Linux systems, network troubleshooting, observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux system administration, Cloud platform management (AWS/Azure/GCP), Python/Bash scripting, Network troubleshooting (TCP/IP, DNS, TLS), HTTP troubleshooting, Observability stack usage
Preferred skills
Kubernetes, Docker, Message queues (Kafka/RabbitMQ), Database troubleshooting (MySQL/PostgreSQL/Redis), ITIL fundamentals
Technologies
Linux (RHEL/Ubuntu), AWS, Azure, GCP, Terraform, Python, Bash, Ansible, tcpdump, Wireshark, Prometheus, Grafana, Datadog, ELK, Nginx, HAProxy
Responsibilities
Own L2/L3 incident response and post-mortems, Monitor system health and SLA/SLO adherence, Automate operational tasks, Collaborate on deployment reliability and capacity planning, Participate in on-call rotation and maintain runbooks
Seniority
Mid-level (3–8 years), hands-on IC