CareerPlanSign in

Senior Site Reliability Engineer

Germany🌐 Remote💼 Full-time🗓 2026-07-20 → 2026-09-25

Core

Ensure reliability, performance, and resilience of high-performance infrastructure and AI inference platforms serving developers and businesses.

Role type

Senior Site Reliability Engineer (distributed systems & GPU workloads)

Builds

Production-grade GPU infrastructure, serverless AI model deployment platforms, and scalable inference services

Domain

Cloud infrastructure, AI/ML workloads, distributed systems

Deliverable

production ML models | infrastructure

Required skills

distributed systems debugging, observability (metrics/logs/tracing), SRE principles (SLIs/SLOs/error budgets), Kubernetes, containers, IaC, automation scripting (Python/Go/PHP), incident management, capacity planning

Preferred skills

high-throughput/low-latency API operations, bare-metal infrastructure, GPU environments, AI/ML workloads, RabbitMQ, MySQL/Redis/ClickHouse, global traffic management, automated scaling/self-healing systems

Responsibilities

Own reliability/availability/performance of critical production services; define and evolve reliability practices (SLIs, SLOs, alerting); investigate complex production issues across distributed systems and GPU workloads; lead incident reviews and RCAs; reduce operational toil via automation and deployment safety improvements; collaborate on capacity planning and architectural scaling

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.