Senior Site Reliability Engineer
Core
Build and operate a Kubernetes-based internal platform hosting AI-driven workflows, evolving from traditional ops to a PaaS model.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Internal developer platform (IDP) and Kubernetes infrastructure for AI/ML workflows
Domain
Cloud Infrastructure / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, AWS, Terraform, CI/CD, Observability, Incident Management, Mentoring
Preferred skills
Internal Developer Platform (IDP) concepts, AI agent orchestration, Multi-cloud, Service mesh, GitOps
Technologies
Kubernetes, AWS, Terraform, Grafana, Splunk, APM, Istio, Linkerd, ArgoCD, Flux, GitHub, GitLab
Responsibilities
Cluster management, networking, and security posture; Developing self-service capabilities and automation; Troubleshooting production issues and driving postmortems; Improving CI/CD pipelines and IaC practices; Building observability and monitoring coverage; Mentoring junior engineers
Seniority
Senior, hands-on IC
