Senior Site Reliability Engineer
Core
Design, build, and maintain core infrastructure and tooling for engineering teams while leading observability practices and incident response for a cloud-native, AI-native platform.
Role type
Senior Site Reliability Engineer (Observability & Infrastructure)
Builds
Cloud-native, no-code platform with embedded AI for industrial operations
Domain
Industrial IoT / Manufacturing Execution Systems (MES) / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Observability tools (Loki, Grafana, Tempo, Mimir), OpenTelemetry, Prometheus, time-series data analysis (promQL), distributed systems debugging, Kubernetes, Go, TypeScript, AI process iteration (Claude Skills/Gemini)
Preferred skills
Mentoring on reliability culture, defining SLIs/SLOs, setting technical agenda
Technologies
Kubernetes, MongoDB, Postgres, Grafana, Loki, Mimir, Tempo, Alloy, Prometheus, OpenTelemetry, Go, TypeScript
Responsibilities
Mentor engineering teams on observability best practices and reliability culture; Perform incident response and debug production issues across the entire stack; Design, build, and maintain core infrastructure and tooling for engineering teams; Contribute to and maintain triage and remediation processes
Seniority
Senior, hands-on IC