SRE
Core
Design, build, and improve OpenTelemetry-based telemetry pipelines for metrics, traces, and logs to strengthen the reliability and visibility of production infrastructure.
Role type
Senior Site Reliability Engineer (Observability)
Builds
Production-grade observability systems on Google Cloud and Kubernetes
Domain
Automotive care industry / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
OpenTelemetry pipeline design, Google Kubernetes Engine (GKE), Google Cloud Monitoring, Cloud Trace, Managed Service for Prometheus, Grafana, Terraform, Helm, SLI/SLO definition, incident response, root-cause analysis, Go/Node.js code review
Preferred skills
Autonomous infrastructure decision-making, telemetry architecture tradeoff analysis, application-level instrumentation collaboration
Technologies
OpenTelemetry, Google Cloud Platform (GCP), Kubernetes, GKE, Prometheus, Grafana, Terraform, Helm, Go, Node.js
Responsibilities
Design and build OpenTelemetry pipelines for metrics, traces, and logs; Integrate telemetry with GCP monitoring and tracing tools; Define and implement SLIs and SLOs; Build actionable burn-rate alerts; Manage observability infrastructure via Terraform and Helm; Participate in on-call and incident-response activities; Lead or contribute to root-cause analyses; Collaborate with software engineers on instrumentation changes; Document operational standards
Seniority
Senior, hands-on IC