Site Reliability Engineer for Observability (Freelance)
Core
Design, implement, and embed a scalable observability foundation using OpenTelemetry, defining SLIs/SLOs/SLAs for the Portable Stack and AI domains.
Role type
Interim Site Reliability Engineer (Observability)
Builds
Scalable observability foundation (OpenTelemetry collector layer) for Azure Databricks, AWS Kubernetes, Azure Kubernetes, AWS, and Azure environments.
Domain
Cloud Infrastructure & Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, containerized workloads, Azure, AWS, infrastructure as code, GitOps, OpenTelemetry, Grafana, Prometheus, Loki, Tempo, SLO/SLI/SLA management, incident management, root cause analysis, Python, Bash, Go, CI/CD
Preferred skills
None stated
Technologies
OpenTelemetry, Azure Databricks, AWS Kubernetes, Azure Kubernetes, Grafana, Prometheus, Loki, Tempo
Responsibilities
Set up an OpenTelemetry collector layer for multiple cloud sources; define and implement SLIs, SLOs, and SLAs for specific domains; enable teams to use new observability capabilities.
Seniority
Experienced, hands-on IC