Site Reliability Engineer, AI Platform
Core
Build and operate production infrastructure supporting AI-related workloads, Kubernetes platforms, and cloud services to ensure reliability and efficiency for Algolia's AI ecosystem.
Role type
Senior Site Reliability Engineer (AI Platform)
Builds
Highly available Kubernetes-based platforms, CI/CD pipelines, and production infrastructure for AI services
Domain
Cloud infrastructure, Kubernetes, AI platform operations
Deliverable
production ML models | infrastructure
Required skills
Kubernetes production operations, Infrastructure as Code, CI/CD pipeline automation, major cloud provider expertise (GCP/AWS/Azure), distributed systems knowledge, observability and troubleshooting, automation mindset
Preferred skills
Go or Python engineering, AI/ML infrastructure exposure, experience with AI-first workflows and coding agents
Technologies
Kubernetes, GCP, AWS, Azure, Infrastructure as Code, CI/CD tools, observability platforms
Responsibilities
Operate and improve highly available Kubernetes-based platforms, improve reliability through SLOs and capacity management, investigate production issues and implement durable fixes, improve CI/CD pipelines and deployment automation, participate in on-call and incident response, collaborate with engineers to take ownership of broader production areas
Seniority
Senior, hands-on IC
