Site Reliability Engineer
Core
Improve cloud infrastructure stability, efficiency, and deployment processes across European regions using GitOps and Kubernetes.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
High-availability cloud environments, automated deployment pipelines, and observability stacks for global services.
Domain
Cloud Infrastructure / DevOps / SRE
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (EKS), GitOps (ArgoCD, Helm), Infrastructure-as-Code (Terraform), Cloud Platforms (AWS), Observability (Prometheus, Loki, Tempo, OpenTelemetry), Incident Response, SLI/SLO Definition, Scripting (Bash, Python, Golang), Networking (TCP/IP, HTTP), Linux OS Optimization, Cache Management (Redis, CDN)
Preferred skills
Rust, Service Mesh (Cilium), JVM Optimization, Real User Monitoring (Grafana Faro)
Technologies
AWS (EKS, EC2, VPC, Lambda, S3, CloudWatch), Kubernetes, ArgoCD, Helm, Terraform, Prometheus, Mimir, Grafana, Alertmanager, Loki, Vector, Tempo, OpenTelemetry, Pyroscope, Grafana Faro, Nginx, Kong, Cilium, eBPF, Docker, Jenkins, GitHub Actions, Aurora MySQL, PostgreSQL, MongoDB, Apache RocketMQ, Kafka, ElastiCache, Valkey, Cloudflare, AWS CloudFront
Responsibilities
Maintain and optimize Kubernetes platform stability and resource utilization; Design and manage alert pipelines to prevent alert fatigue; Own weekend on-call operations and lead post-incident reviews; Define and maintain SLIs and SLOs for critical services; Liaise with external security agencies for audits and perform internal security sweeps; Mentor less experienced team members.
Seniority
Senior, hands-on IC