Senior Site Reliability Engineer - Volcano
Core
Senior SRE owning end-to-end reliability for Volcano, an internal developer platform providing preview environments, edge deployments, and managed data services for Kong's engineering ecosystem.
Role type
Senior Staff/Principal Site Reliability Engineer (Internal Developer Platform)
Builds
Multi-region Kubernetes infrastructure, GitOps/CI/CD pipelines, managed PostgreSQL/Redis/Object Storage, and edge deployment capabilities.
Domain
Developer Platform Engineering / Internal Tooling
Deliverable
production ML models | product features | infrastructure
Required skills
SRE practices (SLOs, error budgets, incident response), Kubernetes (multi-tenant, networking, autoscaling), GitOps (ArgoCD, Helm), Infrastructure as Code (Terraform), Observability (Datadog, Prometheus, Grafana), PostgreSQL administration, Disaster Recovery
Preferred skills
Greenfield platform design, Serverless compute, Vector databases, AI-native infrastructure components
Technologies
Kubernetes, ArgoCD, Helm, Terraform, Terragrunt, Datadog, Prometheus, Grafana, PostgreSQL, Redis
Responsibilities
Define and drive SLOs and incident response practices for Volcano services; Design and build multi-region Kubernetes infrastructure and data plane; Establish deployment automation and preview environment provisioning; Design and operate multi-tenant PostgreSQL clusters and object storage; Instrument services with SLIs and build observability dashboards; Collaborate with OCTO and security to bake reliability into architecture; Evaluate emerging technologies for architectural decisions
Seniority
Senior, hands-on IC (Staff/Principal level)
