Site Reliability Operations Engineer
Core
Incident Commander and reliability engineer supporting global internal business operations, diagnosing complex infrastructure issues, and driving automation to reduce toil.
Role type
Senior Site Reliability Operations Engineer (Incident Management)
Builds
Incident response playbooks, SOPs, and automated remediation for enterprise systems
Domain
Enterprise Technology & Infrastructure (Internal Operations)
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident command, root cause analysis, cross-functional coordination, ITIL framework, cloud platforms, monitoring tools, scripting, problem management
Preferred skills
Salesforce platform, ITIL/AWS/CCNA certifications, Python/Bash/PowerShell, Splunk/Grafana, Puppet/Chef
Technologies
AWS, Linux, Windows, Splunk, Grafana, Tableau, Puppet, Chef, Python, Bash, PowerShell
Responsibilities
Respond to and manage major incidents as Incident Commander; monitor and troubleshoot enterprise systems; create and improve runbooks and SOPs; coordinate emergency changes; analyze incident data and KPIs; lead problem management activities; participate in on-call rotation as Duty Manager
Seniority
Senior, hands-on IC