Site Reliability Engineer - CTJ - Poly
Core
Architect, implement, and optimize hybrid and cloud infrastructure; design data governance and disaster recovery for a multi-petabyte Azure environment; build large-scale data pipelines and automation to ensure service resiliency.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure & Data Engineering)
Builds
Hybrid and cloud infrastructure, data pipelines, disaster recovery systems, and automation tooling for mission-critical services
Domain
Cloud Infrastructure (Azure) and Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Infrastructure as Code (Terraform, Bicep), Cloud Architecture (Azure), Data Pipeline Design, Disaster Recovery Planning, On-call Incident Management, Technical Documentation, Automation Scripting
Preferred skills
Container orchestration (Kubernetes, AKS), Big Data processing (Spark, Hadoop), Cloud platforms (AWS, GCP), Programming (C#, Java, Python, PowerShell, Bash)
Technologies
Azure, Terraform, Bicep, AKS, Docker, Kubernetes, Spark, Hadoop, Python, PowerShell, Bash, C#, Java (via careerplan.io/jobs/1970393556868612-site-reliability-engineer-ctj-poly-at-microsoft)
Responsibilities
Architect and optimize hybrid/cloud infrastructure for availability and security; Design and implement data governance, storage, backup, and disaster recovery; Build and operate large-scale data pipelines; Participate in on-call rotation for troubleshooting and incident response; Identify and implement automation to reduce manual toil; Create technical documentation and design specifications.
Seniority
Senior, hands-on IC