Sr. Observability Engineering
Core
Design, deploy, and operate enterprise observability platforms (Dynatrace, Grafana, Prometheus, ThousandEyes) and implement OpenTelemetry standards for cloud-native and AI/LLM workloads to ensure system reliability and visibility.
Role type
Senior IC Observability Engineer (Platform Engineering)
Builds
Enterprise observability platforms, AI/LLM telemetry pipelines, and self-service monitoring tooling for McKesson's healthcare technology systems.
Domain
Healthcare IT, Cloud Infrastructure (AWS/Azure/GCP), Observability, SRE
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
OpenTelemetry, Dynatrace, Grafana, Prometheus, ThousandEyes, Kubernetes, Terraform, Python, Go, Bash, SRE principles, GitOps
Preferred skills
AIOps, eBPF, ServiceNow, SIEM integration
Technologies
Dynatrace, LogicMonitor, Grafana, Prometheus, ThousandEyes, OpenTelemetry, Terraform, Helm, Kubernetes, AWS, Azure, GCP
Responsibilities
Engineer and operate enterprise observability platforms; Lead OpenTelemetry adoption; Manage Cisco ThousandEyes for network visibility; Define observability-as-code practices; Apply SRE principles (SLOs, error budgets, chaos engineering); Build self-service dashboards; Lead post-incident reviews; Design observability strategies for cloud-native and AI/LLM workloads; Mentor engineers on instrumentation techniques.
Seniority
Senior, hands-on IC with mentorship responsibilities