Senior Site Reliability Engineering
Core
Build and extend AI-driven agents and automation systems to triage, classify, and resolve incidents for cloud services, reducing manual investigation time and enabling faster customer resolution.
Role type
Senior Site Reliability Engineer (AI/ML focus on incident management)
Builds
AI-driven triage agents, auto-routing systems, incident summarization tools, and postmortem generation workflows
Domain
Cloud services, incident management, AI/ML automation
Deliverable
production ML models | infrastructure
Required skills
C#, PowerShell, Python, KQL/Kusto, incident management systems, monitoring and observability, cloud services operations, AI agent development, automation scripting
Preferred skills
Experience with ICM, PagerDuty, ServiceNow, Kusto, Geneva, Grafana
Technologies
C#, PowerShell, Python, KQL/Kusto, ICM, PagerDuty, ServiceNow, Kusto, Geneva, Grafana
Responsibilities
Provide on-call coverage for incoming incidents across CDI services; Build and extend AI-driven agents that ingest alerts and correlate with deployments; Develop automation matching incidents to Troubleshooting Guides; Configure and extend ICM routing rules and intelligent classification systems; Build agents for incident summarization, customer communications drafting, and postmortem generation
Seniority
Senior, hands-on IC