Site Reliability Engineer 2
Core
Build AI-driven agentic triage systems and automation to ingest alerts, correlate with deployments, and route incidents for cloud services.
Role type
Senior IC Site Reliability Engineer (AI/ML automation)
Builds
AI-driven incident triage agents, automation matching incidents to Troubleshooting Guides, and intelligent workflows for cloud services.
Domain
Cloud infrastructure, Site Reliability Engineering, AI/ML automation
Deliverable
production ML models
Required skills
C#, PowerShell, Python, KQL/Kusto, incident management systems, monitoring and observability, AI/ML-driven automation, Azure, Power BI, Fabric services, Troubleshooting Guide authoring, SLA management
Preferred skills
LLMs, Copilot extensibility, MCP servers, agentic frameworks, log traversal, telemetry analysis, incident pattern analysis
Technologies
C#, PowerShell, Python, KQL/Kusto, ICM, PagerDuty, ServiceNow, Kusto, Geneva, Grafana, Azure, Power BI, Fabric
Responsibilities
Provide on-call coverage for incoming incidents across CDI services; Build and extend AI-driven agents that ingest ICM alerts and produce initial assessments; Develop automation that matches incoming incidents to relevant Troubleshooting Guides and known issues.
Seniority
Senior, hands-on IC