Site Reliability Engineer
Core
Ensure reliability, availability, and performance of client-facing services through observability, automation, and incident response.
Role type
Site Reliability Engineer
Builds
Client-facing services and observability platforms
Domain
Insurance technology / Cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud service management (Azure), Observability tools (Datadog), Automation scripting (Python/PowerShell), Infrastructure as Code (Terraform/Pulumi/ARM/Bicep), CI/CD pipelines (Azure DevOps), Incident response and root cause analysis, Knowledge transfer and mentoring
Preferred skills
Azure certifications, Containerization and orchestration (Docker/Kubernetes), C# programming
Technologies
Azure, Datadog, Python, PowerShell, Terraform, Pulumi, ARM Templates, Bicep, Azure DevOps, Docker, Kubernetes, C#
Responsibilities
Collaborate with cross-functional teams to ensure service reliability; Maintain and configure observability platforms; Proactively monitor production environments; Design and implement automation processes; Lead incident response and troubleshooting; Contribute to service design and operational management; Participate in on-call rotation
Seniority
Mid-level, hands-on IC