Senior Software Engineer - Live Site Reliability
Core
Provide on-call support for customer-facing services, design automation systems for incident triage, and build tools for incident lifecycle management.
Role type
Senior Software Engineer - Live Site Reliability
Builds
Automation systems for incident triage, tools for incident lifecycle management, and classification/routing systems for incoming incidents.
Domain
Cloud services, distributed systems, and operational monitoring.
Required skills
Incident response and triage, log analysis, telemetry-based diagnostics, automation system development, incident lifecycle management, troubleshooting guide authoring, SLA and escalation process understanding, cross-functional collaboration, on-call rotation management.
Preferred skills
AI/ML techniques, large language models (LLMs), incident pattern analysis.
Technologies
C, C++, C#, Java, JavaScript, Python, telemetry systems, alerting systems, production support platforms.
Responsibilities
Monitor and investigate customer-facing service incidents, design automation for alert correlation, build incident reporting tools, analyze operational metrics to drive improvements, partner with engineering teams to reduce operational overhead, contribute to distributed system design.
Seniority
Senior, hands-on IC