CareerPlanSign in

Senior Site Reliability Engineer

Romania🌐 Remote💼 Full-time🗓 2026-09-21 → 2026-09-25

Core

End-to-end ownership of reliability and operational performance for critical production systems in a high-throughput platform.

Role type

Senior Site Reliability Engineer (IC)

Builds

Secure, resilient, cost-efficient infrastructure and developer-facing tooling

Domain

Cloud infrastructure, distributed systems, observability

Deliverable

production ML models | product features | infrastructure

Required skills

Cloud infrastructure (Kubernetes/EKS, networking, load balancing), Infrastructure as Code (Terraform), Programming (Go, Python), Observability (Datadog, Prometheus, Grafana, OpenTelemetry), Redis/ElastiCache management, Distributed systems failure modes, Incident response, Capacity planning, SLO/SLI/Error budget management, Chaos engineering

Preferred skills

AI tools integration, Progressive delivery practices

Technologies

Terraform, Kubernetes, EKS, Datadog, Prometheus, Grafana, OpenTelemetry, Redis, ElastiCache, Go, Python

Responsibilities

Define and maintain SLIs, SLOs, and error budgets; Lead high-severity incident response and postmortems; Design and operate resilient infrastructure with focus on failure modes; Perform capacity and performance analysis; Build production software and tooling to reduce operational toil; Conduct controlled failure testing and game days; Mentor engineers and partner on production readiness

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.