Lead Site Reliability Engineer - Observability
Core
Lead the evolution of a consistent Observability approach across the full SimCorp One ecosystem, ensuring system visibility, uptime, and performance for financial clients.
Role type
Lead Site Reliability Engineer (Observability)
Builds
Unified visibility across the full stack (LLM-driven agents to backend services) in Azure-based environments.
Domain
FinTech, Cloud Native, Observability, AI/LLM workflows
Deliverable
production ML models | infrastructure
Required skills
Telemetry data collection and analysis, OpenTelemetry, Azure Monitor, Infrastructure as Code (Bicep, ARM, Terraform), Kubernetes, Docker, Distributed tracing, SLO/Error budget management, Root cause analysis, Synthetic monitoring, AI/ML anomaly detection, PowerShell, Bash, SQL, Cosmos DB, Postgres SQL
Preferred skills
LangChain, Celery, OpenAI APIs, Checkly, Playwright, SimCorp Dimension, Salesforce, ITIL practices
Technologies
Microsoft Azure, Application Insights, DataDog, Log Analytics, Azure Anomaly Detector, Kubernetes, Docker, OpenTelemetry, Bicep, ARM, Terraform, LangChain, Celery, Playwright, Checkly
Responsibilities
Deploy and manage instrumentation for applications to gain granular insights into service health; Unify observability tooling across teams ensuring metrics, logs, and traces flow into a central platform; Enable and configure OpenTelemetry-based data collection within Azure Monitor; Work with product development teams to enable structured logging, basic distributed tracing, and core metrics; Support incident response by gathering logs, metrics, and traces to perform root cause analysis; Build tools and automation to eliminate TOIL and improve engineering velocity; Define and manage SLOs and error budgets in partnership with Engineering teams.
Seniority
Senior, hands-on IC with leadership responsibilities