Senior Site Reliability Engineer
Core
Design and implement end-to-end telemetry, alerting, self-healing, and automation capabilities to ensure the stability and performance of Microsoft's operational database systems (Azure DocumentDB/CosmosDB).
Role type
Senior Site Reliability Engineer (IC)
Builds
Operational database services (PostgreSQL-based NoSQL, relational, non-relational) for cloud-native applications
Domain
Cloud infrastructure, distributed systems, database operations
Deliverable
production ML models | infrastructure
Required skills
automation/scripting (PowerShell, Python), programming (C++, C#), troubleshooting/debugging, telemetry-based analysis (KQL), distributed systems debugging, network troubleshooting, hardware troubleshooting, code optimization
Preferred skills
distributed systems architecture, networking expertise
Technologies
Azure DocumentDB, Azure CosmosDB, PostgreSQL, KQL, PowerShell, Python, C++, C#
Responsibilities
Design and implement telemetry, alerting, self-healing, and automation; participate in on-call rotations to triage and resolve service issues; interact with customers and product teams to evolve services; own availability, performance, and supportability targets; author functional and technical documentation (via careerplan.io/jobs/1970393556751926-senior-site-reliability-engineer-at-microsoft)
Seniority
Senior, hands-on IC
