Senior Platform Engineer - Observability
Core
Design, build, and operate unified observability capabilities (metrics, events, logs, traces) for thousands of engineers to troubleshoot application and infrastructure health.
Role type
Senior Platform Engineer (Observability)
Builds
Unified observability platform, telemetry pipelines, Observability-as-Code modules, and single-pane-of-glass dashboards.
Domain
Cloud-native infrastructure, SRE, Observability, OpenTelemetry
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python or Go, Infrastructure-as-Code (Terraform/OpenTofu), OpenTelemetry standards, cloud-native architectures (AWS EKS/Kubernetes), telemetry pipeline engineering, AIOps/anomaly detection, SLI/SLO definition, distributed tracing, cost/cardinality optimization
Preferred skills
AI-assisted coding workflows, mentoring engineering teams, vendor-neutral architecture design
Technologies
Datadog, Prometheus, Grafana, Splunk, CloudWatch, OpenTelemetry Collector, GitHub Copilot, Terraform, Kubernetes, AWS
Responsibilities
Engineer telemetry collection and pipelines; automate onboarding for out-of-the-box observability; deliver Observability-as-Code (monitors, dashboards, alerts); correlate MELT data into a unified experience; define golden signals and reliability standards; optimize telemetry cost and cardinality; develop platform tooling using AI-assisted workflows.
Seniority
Senior, hands-on IC with mentorship responsibilities