Senior Site Reliability Engineer
Core
Champion reliability, availability, and performance of an Azure-based healthcare platform by defining SRE practices, managing incidents, and building observability frameworks.
Role type
Senior Site Reliability Engineer (IC)
Builds
Resilient, compliant, and observable cloud infrastructure for a national specialty care platform
Domain
Healthcare technology / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE practices, incident management, observability, Infrastructure-as-Code, CI/CD, scripting, capacity planning, regulatory compliance
Preferred skills
chaos engineering, Azure Kubernetes Service, relevant certifications
Technologies
Azure, Terraform, Datadog, Azure Monitor, Rootly, Python, Bash, PowerShell, Azure DevOps, GitHub Actions
Responsibilities
Define and track SLOs/SLIs/error budgets; Build and maintain observability platforms; Lead incident management processes; Automate operational toil through IaC; Design and implement disaster recovery strategies; Collaborate on architecture reviews and chaos engineering; Optimize system performance and cost; Ensure compliance with HIPAA/SOC 2; Maintain CI/CD pipelines; Mentor junior engineers
Seniority
Senior, hands-on IC