Lead Site Reliability Engineer
Core
Lead the development of SRE solutions, including monitoring, alerting, machine learning-based anomaly detection, self-healing mechanisms, and reliability testing strategies for SimCorp's financial technology products.
Role type
Senior IC Site Reliability Engineer (Azure Cloud)
Builds
Production-grade Azure infrastructure, automation tools, and reliability frameworks for investment management solutions.
Domain
FinTech / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Azure Cloud architecture, Infrastructure as Code (Bicep, ARM, Terraform), Incident management, SLO/Error budget management, Kubernetes, Docker, PowerShell, Bash, SQL, Observability (OpenTelemetry, distributed tracing), Synthetic monitoring, ITIL practices, Mentoring
Preferred skills
SimCorp Dimension, Salesforce, AI/ML anomaly detection tools, Playwright
Technologies
Microsoft Azure, Bicep, ARM, Terraform, Azure Monitor, Application Insights, DataDog, Log Analytics, Open Telemetry, Kubernetes, Docker, Playwright, SQL, Cosmos DB, PostgreSQL
Responsibilities
Lead SRE solution development (monitoring, alerting, anomaly detection); Design reliability and scalability strategies; Build automation tools to reduce TOIL; Define and manage SLOs and error budgets; Manage incident response and root cause analysis; Mentor junior SREs; Participate in on-call rotations.
Seniority
Senior, hands-on IC with leadership responsibilities