CareerPlanSign in

Site Reliability Engineer - CTJ - Poly

United States, Washington, Redmond💼 Full-time🗓 2026-06-24 → 2026-10-02

Core

Architect, implement, and optimize hybrid and cloud infrastructure; design data governance and disaster recovery for a multi-petabyte Azure environment; build large-scale data pipelines and automation to ensure service resiliency.

Role type

Senior Site Reliability Engineer (Cloud Infrastructure & Data Engineering)

Builds

Hybrid and cloud infrastructure, data pipelines, disaster recovery systems, and automation tooling for mission-critical services

Domain

Cloud Infrastructure (Azure) and Data Engineering

Deliverable

production ML models | infrastructure

Required skills

Infrastructure as Code (Terraform, Bicep), Cloud Architecture (Azure), Data Pipeline Design, Disaster Recovery Planning, On-call Incident Management, Technical Documentation, Automation Scripting

Preferred skills

Container orchestration (Kubernetes, AKS), Big Data processing (Spark, Hadoop), Cloud platforms (AWS, GCP), Programming (C#, Java, Python, PowerShell, Bash)

Technologies

Azure, Terraform, Bicep, AKS, Docker, Kubernetes, Spark, Hadoop, Python, PowerShell, Bash, C#, Java (via careerplan.io/jobs/1970393556868612-site-reliability-engineer-ctj-poly-at-microsoft)

Responsibilities

Architect and optimize hybrid/cloud infrastructure for availability and security; Design and implement data governance, storage, backup, and disaster recovery; Build and operate large-scale data pipelines; Participate in on-call rotation for troubleshooting and incident response; Identify and implement automation to reduce manual toil; Create technical documentation and design specifications.

Seniority

Senior, hands-on IC