Reliability Operations Engineer
Core
Ensure the stability, reliability, and continuous improvement of a full-stack cloud communication platform through incident management, monitoring, and automation.
Role type
Platform / Site Reliability Engineer
Builds
Global cloud communication platform
Domain
Cloud communications / SRE
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
incident management, monitoring and alerting, scripting and automation, root cause analysis, impact assessment, technical troubleshooting, cross-functional collaboration
Preferred skills
mentoring, autonomous problem solving, proactive reliability focus
Technologies
(none explicitly listed)
Responsibilities
Create and improve platform alerts and runbooks, monitor platform and triage incidents, execute mitigation actions to restore service, write and maintain automation scripts, mentor other engineers
Seniority
Individual Contributor, operational excellence