Senior Site Reliability Engineer (SRE, Compute Node Team)
Core
Building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions, focusing on Linux systems engineering, virtualization, and operational reliability.
Role type
Senior Site Reliability Engineer (Compute Node)
Builds
Cluster scheduler, node-level services, and virtual machine management infrastructure for a full-stack AI cloud platform
Domain
Cloud Infrastructure / AI Compute / Systems Engineering
Deliverable
production ML models | infrastructure
Required skills
Linux kernel space and user space expertise, QEMU/KVM virtualization, containerization (namespaces, cgroups), complex system debugging (CPU, memory, NUMA, scheduling), observability stack design (metrics, logs, traces, SLIs/SLOs), incident response and root-cause analysis
Preferred skills
Kubernetes internals, low-level Linux debugging tools (perf, eBPF, ftrace, strace), large-scale compute or bare-metal platforms, open-source infrastructure contributions, hardware/driver-level debugging (GPUs, NVLink, InfiniBand)
Technologies
Linux, QEMU, KVM, cgroups, namespaces, Kubernetes, eBPF, perf, ftrace, strace
Responsibilities
Ensure reliability, availability, and performance of compute nodes running VMs; Analyze and debug Linux systems across user space and kernel space; Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling; Work hands-on with virtualization and containerization; Design and evolve observability capabilities; Lead incident response, root-cause analysis, and postmortems; Collaborate with platform, kernel/hypervisor, GPU, and infrastructure teams
Seniority
Senior, hands-on IC