Site Reliability Engineer - CTJ - POLY
Core
Owns reliability architecture and end-to-end service understanding for distributed systems at scale, defining service health via SLIs/SLOs and managing incidents.
Role type
Senior Site Reliability Engineer
Builds
Distributed systems, cloud platforms, and operational tooling (via careerplan.io/jobs/1970393556761670-site-reliability-engineer-ctj-poly-at-microsoft)
Domain
Cloud infrastructure and distributed systems
Required skills
Linux administration, automation (Ansible), CI/CD pipeline development, incident response, observability (metrics/logs/traces), SLO/SLI definition, infrastructure-as-code, secure-by-design operations
Preferred skills
Chaos engineering, progressive delivery, feature flags, blameless postmortems, security compliance integration
Technologies
Ansible, Azure DevOps, GitHub Actions, Rocky 9, Redhat, Mariner
Responsibilities
Define and improve service health via SLIs/SLOs and error budgets; lead response for complex, high-impact incidents; build automation to reduce toil and enable self-healing; expand infrastructure-as-code and operational tooling; produce blameless postmortems with corrective actions; implement secure-by-design operations and compliance controls
Seniority
Senior, hands-on IC