CareerPlanGet AI match score →

Senior Site Reliability Engineer (SRE, Compute Node Team)

Amsterdam💼 Full-time🗓 2026-07-01 → 2026-07-31

Core

Building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions, focusing on Linux systems engineering, virtualization, and operational reliability.

Role type

Senior Site Reliability Engineer (Compute Node)

Builds

Cluster scheduler, node-level services, and virtual machine management infrastructure for a full-stack AI cloud platform

Domain

Cloud Infrastructure / AI Compute / Systems Engineering

Deliverable

production ML models | infrastructure

Required skills

Linux kernel space and user space expertise, QEMU/KVM virtualization, containerization (namespaces, cgroups), complex system debugging (CPU, memory, NUMA, scheduling), observability stack design (metrics, logs, traces, SLIs/SLOs), incident response and root-cause analysis

Preferred skills

Kubernetes internals, low-level Linux debugging tools (perf, eBPF, ftrace, strace), large-scale compute or bare-metal platforms, open-source infrastructure contributions, hardware/driver-level debugging (GPUs, NVLink, InfiniBand)

Technologies

Linux, QEMU, KVM, cgroups, namespaces, Kubernetes, eBPF, perf, ftrace, strace

Responsibilities

Ensure reliability, availability, and performance of compute nodes running VMs; Analyze and debug Linux systems across user space and kernel space; Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling; Work hands-on with virtualization and containerization; Design and evolve observability capabilities; Lead incident response, root-cause analysis, and postmortems; Collaborate with platform, kernel/hypervisor, GPU, and infrastructure teams

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗