Staff Site Reliability Engineer
Core
Staff AI-SRE managing large-scale, global production environments, pioneering AI-driven automation and intelligent infrastructure operations.
Role type
Staff AI-Site Reliability Engineer
Builds
Self-healing infrastructure, AI-powered operational runbooks, predictive scalability systems, and intelligent alerting/diagnostics.
Domain
Cybersecurity, Cloud Infrastructure, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
Linux system internals, Kubernetes (GKE), API Gateway (Kong), Terraform, Python, AIOps platforms, LLM integration, incident management, security compliance
Preferred skills
Service Mesh (Istio, Kong Mesh), CI/CD tools (GitLab CI, ArgoCD), AWS/Azure experience, CKA/CKAD certifications
Technologies
Kubernetes, GKE, Kong, Terraform, Helm, Python, Bash, Go, LLMs, GitHub Copilot, Claude
Responsibilities
Operate large-scale global production environments with AI focus; Pioneer AI-driven automation (LLM runbooks, intelligent alerting); Design self-healing infrastructure; Lead Kong API Gateway architecture and operations; Manage production-grade GKE clusters; Implement predictive scalability using AI modeling; Resolve P1/P2 incidents using AI-assisted diagnostics; Drive end-to-end troubleshooting with AI tools; Implement Infrastructure as Code with AI integration; Champion AI-first workflows and tooling development.
Seniority
Staff, hands-on IC with strategic leadership