Staff Observability Platform Engineer (SRE)
Core
Design and implement metrics, observability frameworks, and automation for cloud infrastructure reliability, focusing on SLOs, error budgets, and AI-driven incident response.
Role type
Staff Observability Platform Engineer (SRE)
Builds
Cloud-based monitoring solutions, automated quality gates, and AI-enhanced observability pipelines for CVS Health PBM.
Domain
Health care / Cloud-native observability / AIOps
Deliverable
production ML models | infrastructure
Required skills
SLO/SLI/SLA design, OpenTelemetry instrumentation, Kubernetes/Argo CD, AWS/GCP/Azure, distributed data pipelines, Grafana/Loki/Prometheus, relational databases, CI/CD automation, GenAI for incident triage, LLM observability governance
Preferred skills
Service meshes (Envoy/Istio), commercial observability platforms (Splunk/AppDynamics), streaming platforms (Kafka/Pulsar), time-series/NoSQL databases, Infrastructure as Code (Terraform), chaos engineering, security-aware platform design, mentoring senior engineers
Technologies
Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Docker, Kubernetes, Argo CD, AWS, GCP, Azure, PostgreSQL, MySQL, Kafka, Pulsar, ClickHouse, Bigtable, Cassandra, Terraform, CloudFormation, Envoy, Istio, Splunk, AppDynamics
Responsibilities
Define and maintain key performance metrics, SLOs, and SLIs; manage error budgets and analyze incidents; design real-time monitoring solutions; architect scalable cloud infrastructure; develop automated quality gates for CI/CD; assist in incident response and post-mortem analyses; implement AI-driven insights for anomaly detection and action surfacing; monitor AI workloads for quality, safety, and latency; embed AI signal checks into release pipelines
Seniority
Staff, hands-on IC with strategic influence