CareerPlanGet AI match score →

Senior Site Reliability Engineer — Token Factory (Inference Platform)

Amsterdam💼 Full-time🗓 2026-07-01 → 2026-07-31

Core

Own the reliability, performance, and observability of a massive GPU inference platform serving foundation models (text, vision, audio) at scale.

Role type

Senior Site Reliability Engineer (Inference Platform)

Builds

High-throughput inference stack for multimodal AI models

Domain

Cloud Infrastructure / AI Inference / GPU Orchestration

Deliverable

production ML models

Required skills

Kubernetes, Prometheus, Grafana, Terraform, Infrastructure-as-Code, Python, Bash, Alert Design, SLOs, Distributed Systems Debugging, GPU Workload Management

Preferred skills

vLLM, Triton, Ray, MLOps, Model Hosting Platforms, Kernel-level Debugging

Technologies

Kubernetes, Prometheus, Grafana, Terraform, vLLM, Triton, Ray

Responsibilities

Design and refine telemetry pipelines for metrics, logs, and traces; tune Kubernetes autoscalers for GPU efficiency; craft Terraform modules for cluster resilience; harden request-routing and retry logic; detect, isolate, and remediate incidents via automation; drive post-mortem culture.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗