CareerPlanSign in

Staff Site Reliability Engineer — AI Platform

Amsterdam, Netherlands💼 Full-time🗓 2026-09-03 → 2026-09-29

Core

Own reliability, performance, and cost of AI infrastructure (inference services, AI Gateway) for a global conversational growth platform.

Role type

Staff Site Reliability Engineer (AI Platform)

Builds

AI Gateway, inference services, integrations with LLM providers (Bedrock, Azure OpenAI, etc.)

Domain

AI/ML Infrastructure, Cloud-Native Systems

Deliverable

production ML models

Required skills

SRE/Platform engineering at scale, LLM system operations, Cloud-native (AWS, Kubernetes, Terraform), Observability (SLOs, token metrics), Cost optimization (FinOps), Technical leadership

Preferred skills

LLM gateway/proxy experience, GPU workload optimization, Eval pipelines for LLM outputs

Technologies

AWS, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry, Amazon Bedrock, Azure OpenAI, vLLM, TGI, Triton, LiteLLM, Kong AI Gateway

Responsibilities

Design and evolve the AI Gateway (routing, failover, rate limiting), Build observability for AI systems (latency, throughput, quality signals), Drive cost optimization and FinOps, Run capacity planning and incident response, Scale AI expertise across the org

Seniority

Staff, hands-on IC with strategic influence

Sourced via ashby · Listed on CareerPlan, which tracks 849,000+ jobs from 20+ sources.