Software Engineer II
Core
Identify systemic health issues, drive recovery programs, and improve service reliability through automation and self-healing capabilities for a large-scale storage fleet.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Automated recovery workflows, monitoring solutions, dashboards, and health signals for cloud storage infrastructure.
Domain
Cloud infrastructure, large-scale storage, datacentre hardware, and distributed systems.
Deliverable
production ML models | infrastructure
Required skills
Root-cause analysis, automation design, cloud architecture, observability, incident management, large-scale system design, server hardware knowledge, cloud environment management.
Preferred skills
AI tools integration, self-healing system design, fleet-wide health indicator analysis.
Technologies
C#, C++, Python, Bash, Powershell, Azure, AWS, GCP.
Responsibilities
Analyse telemetry and incident trends to identify automation opportunities; Lead investigations of production incidents and drive root-cause analysis; Design and implement safe recovery workflows; Create dashboards and monitoring solutions; Act as a Designated Responsible Individual (DRI) to guide engineers and restore services during outages.
Seniority
Senior, hands-on IC