Site Reliability Engineer - Observability
Core
Design, deploy, and operate observability pipelines for logs, metrics, traces, and alerts across Proton's services using open-source technologies.
Role type
Senior Site Reliability Engineer (Observability)
Builds
Observability infrastructure (logs, metrics, traces, alerts) and AI-powered detection tooling for Proton's encrypted services
Domain
Privacy-focused SaaS / Observability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, Go, Kubernetes, GitOps, Infrastructure-as-Code (Terraform, Ansible, Puppet), Linux system administration, Open-source observability stacks
Preferred skills
OpenTelemetry, ClickHouse, AI/ML tooling
Technologies
Python, Go, Kubernetes, ArgoCD, Terraform, Ansible, Puppet, ClickHouse, OpenTelemetry, Linux
Responsibilities
Design and operate observability pipelines; Partner with teams to ship alerting and dashboarding solutions; Build reusable tooling for incident response; Champion observability best practices; Develop AI-powered detection capabilities; Evolve the platform iteratively
Seniority
Senior, hands-on IC