Staff Engineer – Observability Platform
Core
Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics to drive platform-wide improvements and uncover the truth behind complex system behaviors.
Role type
Staff Engineer (Observability Platform)
Builds
World-class observability capabilities, tooling, and automation for incident detection and diagnosis.
Domain
Cloud-native distributed systems, observability, and platform engineering.
Deliverable
production ML models | product features | infrastructure
Required skills
distributed systems architecture, observability domains (monitoring, logging, tracing, telemetry), system debugging and root-cause analysis, Go/Java/Python/Node.js programming, influence without authority, technical vision and architecture.
Preferred skills
internal developer platforms, AIOps and anomaly detection, reliability initiatives across multiple organizations, SRE/Platform Engineering background.
Technologies
OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb.
Responsibilities
Drive technical vision and architecture for the observability platform; Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics; Investigate complex production issues and drive long-term corrective actions; Partner with engineering teams to improve service reliability, availability, performance, and operational maturity; Establish observability standards, best practices, and instrumentation frameworks; Mentor senior engineers and raise the technical bar across the organization.
Seniority
Staff, hands-on IC with significant mentorship and cross-functional influence.