CareerPlanSign in

Lead Site Reliability Engineer - Observability

Hyderabad💼 Full-time🗓 2026-06-24 → 2026-09-25

Core

Lead the evolution of a consistent Observability approach across the full SimCorp One ecosystem, ensuring system visibility, uptime, and performance for financial clients.

Role type

Lead Site Reliability Engineer (Observability)

Builds

Unified visibility across the full stack (LLM-driven agents to backend services) in Azure-based environments.

Domain

FinTech, Cloud Native, Observability, AI/LLM workflows

Deliverable

production ML models | infrastructure

Required skills

Telemetry data collection and analysis, OpenTelemetry, Azure Monitor, Infrastructure as Code (Bicep, ARM, Terraform), Kubernetes, Docker, Distributed tracing, SLO/Error budget management, Root cause analysis, Synthetic monitoring, AI/ML anomaly detection, PowerShell, Bash, SQL, Cosmos DB, Postgres SQL

Preferred skills

LangChain, Celery, OpenAI APIs, Checkly, Playwright, SimCorp Dimension, Salesforce, ITIL practices

Technologies

Microsoft Azure, Application Insights, DataDog, Log Analytics, Azure Anomaly Detector, Kubernetes, Docker, OpenTelemetry, Bicep, ARM, Terraform, LangChain, Celery, Playwright, Checkly

Responsibilities

Deploy and manage instrumentation for applications to gain granular insights into service health; Unify observability tooling across teams ensuring metrics, logs, and traces flow into a central platform; Enable and configure OpenTelemetry-based data collection within Azure Monitor; Work with product development teams to enable structured logging, basic distributed tracing, and core metrics; Support incident response by gathering logs, metrics, and traces to perform root cause analysis; Build tools and automation to eliminate TOIL and improve engineering velocity; Define and manage SLOs and error budgets in partnership with Engineering teams.

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.