Staff Observability Engineer
Core
Define and implement observability strategy, frameworks, and self-remediation solutions for cloud-native applications to ensure system health and resilience.
Role type
Staff Software Engineer (Observability)
Builds
Standardized observability frameworks, dashboards, alerts, SLOs, and self-remediation solutions for PCS DS applications.
Domain
Cloud-native software engineering, observability, healthcare compliance
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Observability pillars (metrics, logs, traces), OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, Go, Python, Bash, Terraform, distributed tracing, SLO/SLI frameworks, incident response workflows, distributed systems, microservices, cloud platforms (AWS, Azure, GCP), AI-powered anomaly detection, chaos engineering
Preferred skills
Healthcare industry experience, data privacy and compliance (HIPAA, HITRUST), cost optimization, telemetry data governance, open-source observability contributions
Technologies
Kubernetes, serverless, OpenTelemetry, Prometheus, Fluent bit, Grafana, Datadog, Dynatrace, AWS, Azure, GCP
Responsibilities
Define and evolve observability vision and roadmap, design and implement standardized observability frameworks, collaborate with teams to instrument services, build and maintain dashboards and alerts, evaluate and optimize observability agents, design self-remediation solutions, lead incident analysis and postmortem reviews, conduct Operational Readiness Reviews, ensure compliance with healthcare standards, mentor engineers
Seniority
Staff, hands-on IC with strategic leadership