Staff Site Reliability Engineer — AI Platform
Core
Own reliability, performance, and cost of AI infrastructure (inference services, AI Gateway) for a global conversational growth platform.
Role type
Staff Site Reliability Engineer (AI Platform)
Builds
AI Gateway, inference services, integrations with LLM providers (Bedrock, Azure OpenAI, etc.)
Domain
AI/ML Infrastructure, Cloud-Native Systems
Deliverable
production ML models
Required skills
SRE/Platform engineering at scale, LLM system operations, Cloud-native (AWS, Kubernetes, Terraform), Observability (SLOs, token metrics), Cost optimization (FinOps), Technical leadership
Preferred skills
LLM gateway/proxy experience, GPU workload optimization, Eval pipelines for LLM outputs
Technologies
AWS, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry, Amazon Bedrock, Azure OpenAI, vLLM, TGI, Triton, LiteLLM, Kong AI Gateway
Responsibilities
Design and evolve the AI Gateway (routing, failover, rate limiting), Build observability for AI systems (latency, throughput, quality signals), Drive cost optimization and FinOps, Run capacity planning and incident response, Scale AI expertise across the org
Seniority
Staff, hands-on IC with strategic influence