Senior Observability Engineer
Core
Lead the development and scaling of telemetry systems to keep platforms reliable, performant, and resilient, shaping observability practices across engineering and infrastructure teams.
Role type
Senior Observability Engineer
Builds
Scalable, automated telemetry pipelines and observability solutions for cloud-native, distributed systems.
Domain
Cloud-native infrastructure, distributed systems, SRE
Deliverable
infrastructure
Required skills
Telemetry design, metrics/logs/traces/events, Terraform, Kubernetes, service mesh architectures, distributed systems behavior, SLOs, error budgets, incident response, root cause analysis
Preferred skills
AI-assisted development workflows, cross-functional collaboration, technical guidance, culture of reliability
Technologies
Splunk Observability Cloud, Google Cloud Observability, Terraform, Kubernetes
Responsibilities
Lead architecture and delivery of observability solutions; Build automated telemetry pipelines with IaC; Define best practices for metrics, traces, logs, and events; Collaborate with app/platform/security teams to integrate observability; Drive adoption of SLIs and error budgets; Lead adoption of AI-assisted debugging workflows; Participate in on-call rotations; Provide technical guidance to engineers; Champion reliability through enablement.
Seniority
Senior, hands-on IC