Observability Architect
Core
Define the strategic vision, technical architecture, and engineering standards for observability across the organization's cloud platforms to enable scalable, cost-efficient, and highly reliable insight into distributed systems.
Role type
Senior IC Observability Architect (SRE)
Builds
Enterprise-wide observability platforms and data pipelines for distributed systems
Domain
Cloud infrastructure, SRE, distributed systems monitoring
Deliverable
production ML models | infrastructure
Required skills
Enterprise-scale observability platform design, OpenTelemetry ecosystem mastery, Prometheus-compatible metrics systems, tracing systems, log aggregation platforms, Terraform, Helm charts, SLO/SLI frameworks, multi-cloud infrastructure provisioning
Preferred skills
None stated
Technologies
Grafana, Prometheus, VictoriaMetrics, Tempo, Loki, Elastic Stack, OpenTelemetry, Terraform, Helm, Google BigQuery, Jaeger, Elasticsearch
Responsibilities
Define and own the enterprise-wide observability architecture and multi-year roadmaps; Evaluate, select, and standardize observability tooling to reduce sprawl; Design scalable data pipelines for petabyte-scale telemetry; Design Terraform modules and Helm charts for multi-cloud provisioning; Establish instrumentation standards using OpenTelemetry; Define and champion SLO/SLI/error-budget frameworks; Serve as senior escalation point during critical incidents; Provide architectural mentorship to Observability Engineers and SREs
Seniority
Senior, hands-on IC with strategic scope