Senior Site Reliability Engineer, AI Platform
Core
Design and operate highly available Kubernetes-based platforms supporting AI-related workloads, model serving, and production AI ecosystems.
Role type
Senior Site Reliability Engineer (AI Platform)
Builds
Shared production foundations for Algolia's AI ecosystem, including Kubernetes clusters, cloud infrastructure, and CI/CD pipelines.
Domain
Cloud Infrastructure, Site Reliability Engineering, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, cloud-native systems, distributed systems, networking, reliability engineering, SLOs, observability, capacity planning, FinOps, automation, incident response
Preferred skills
Go, Python, AI/ML workload support, GPU compute, agentic development workflows, AI-assisted debugging
Technologies
Kubernetes, GCP, AWS, Azure
Responsibilities
Own and evolve production infrastructure for AI workloads; drive reliability through SLOs and guardrails; lead complex production investigations; improve shared infrastructure (networking, databases, compute); build CI/CD and progressive delivery; drive FinOps initiatives; participate in on-call and incident response; mentor engineers.
Seniority
Senior, hands-on IC with cross-team leadership
