Senior Lead Engineer, Platform Operations & Observability
Core
Lead reliability, observability, and operational excellence for enterprise healthcare technology platforms, focusing on monitoring strategies, incident management, and production readiness.
Role type
Senior Lead Engineer, Platform Operations & Observability
Builds
Scalable, secure, and resilient enterprise healthcare platforms (B2C Pharmacy Patient Platform)
Domain
Healthcare / Enterprise Software / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident management, Root cause analysis (RCA), Observability strategy, CI/CD pipeline leadership, Microservices architecture, Cloud platform management, Change governance, Production readiness reviews, Technical mentorship
Preferred skills
Kubernetes, Container orchestration, Azure cloud, ITIL-aligned practices, SLA/MTTR definition, Platform modernization, AI tools (Copilot, Rovo)
Technologies
Dynatrace, Prometheus, Dotcom Monitor, Kubernetes, Azure
Responsibilities
Lead monitoring and observability strategies across enterprise applications; Design and optimize dashboards, alerts, telemetry, and logging solutions; Drive incident management processes and postmortem analysis; Partner with engineering to improve platform reliability and scalability; Lead change management reviews and safe deployment practices; Provide technical leadership and coaching to engineering teams
Seniority
Senior, hands-on IC with leadership responsibilities