Senior Site Reliability Engineer
Core
Design and maintain a Kubernetes-based multi-cloud platform, improve observability tools, and drive automation to ensure reliable, scalable infrastructure for developers.
Role type
Senior Site Reliability Engineer (IC)
Builds
Kubernetes clusters, monitoring/alerting systems, automation scripts, and runbooks
Domain
Cloud Infrastructure / DevOps / AI Platform
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Infrastructure as Code (Terraform), Monitoring/Observability (Prometheus/Grafana), Incident Response, Root Cause Analysis, Automation
Preferred skills
AWS/GCP managed Kubernetes, GitOps (ArgoCD), Python/Go, SLOs
Technologies
Kubernetes, Terraform, Prometheus, Grafana, AWS, Google Cloud Platform, ArgoCD, Python, Go
Responsibilities
Design and evolve multi-cloud platform architecture; Implement and improve monitoring and alerting tools; Participate in on-call rotations and resolve production incidents; Collaborate with product teams to define and deliver infrastructure features; Automate repetitive tasks and share knowledge; Mentor less experienced engineers on infrastructure challenges
Seniority
Senior, hands-on IC
