Lead Site Reliability Engineer
Core
Lead SRE for Azure-based SaaS environments, ensuring reliability, scalability, and automation for client onboarding and production operations.
Role type
Lead Site Reliability Engineer (Cloud Infrastructure)
Builds
Azure-hosted SaaS platforms, automated onboarding pipelines, observability frameworks
Domain
FinTech, Cloud Infrastructure (Azure)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Azure production design, IaC (Terraform, Bicep, ARM, Ansible), cloud-native monitoring, incident management, Kubernetes, Docker, Linux/Windows systems, scripting (PowerShell, Bash), SQL
Preferred skills
SimCorp Dimension experience, networking, virtualization, API design
Technologies
Azure, Terraform, Bicep, ARM, Ansible, Kubernetes, Docker, Azure Monitor, Application Insights, Log Analytics, Grafana, PowerShell, Bash, SQL
Responsibilities
Own reliability and scalability of Azure environments; lead operational support and client onboarding; provide technical leadership in SRE practices and incident response; automate manual processes; lead solutions workshops and manage senior stakeholders; implement observability frameworks with SLOs/SLIs; oversee disaster recovery and root cause analysis; mentor engineers; coordinate maintenance, deployments, and rollbacks; collaborate on platform design and roadmaps.
Seniority
Senior, hands-on IC with leadership responsibilities