Core
Hands-on Site Reliability Engineer focused on cloud operations, deployment reliability, infrastructure automation, and production support for customer-facing SaaS and data platform environments.
Role type
Senior IC Site Reliability Engineer
Builds
Stable, repeatable deployment and support processes for Azure-hosted SaaS and data platform production environments
Domain
Cloud Infrastructure (Azure) & DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Azure Kubernetes Service (AKS), Terraform, Infrastructure as Code (IaC), CI/CD pipelines, Kubernetes operations, GitOps, observability, incident response, release automation, secret management, RBAC, capacity planning
Preferred skills
Azure Data Factory, Azure Databricks, Helm, Flux, Argo CD, Bicep, ARM templates, service principal management
Technologies
Azure, Kubernetes, Terraform, GitHub Actions, Azure DevOps, Helm, Flux, Argo CD, Bicep, ARM
Responsibilities
Support availability, latency, performance, and reliability for production environments; Troubleshoot and resolve operational issues impacting service uptime; Maintain site stability and uptime for 24x7 SaaS platforms; Support automated release, hotfix, and customer deployment processes; Implement and maintain infrastructure as code using Terraform and related tools; Manage Kubernetes clusters, containerized workloads, and GitOps deployment patterns; Assist with identity, access control, and CI/CD secret management; Perform incident root cause analysis and service restoration; Develop and maintain technical documentation including runbooks and deployment guides; Optimize system performance and identify reliability improvements based on monitoring data.
Seniority
Senior, hands-on IC