Senior Site Reliability Engineer
Core
Design, build, and maintain highly scalable and reliable cloud-native infrastructure for a globally distributed, highly available system.
Role type
Senior Site Reliability Engineer (IC)
Builds
Production-grade cloud infrastructure, Kubernetes clusters, and automation tools
Domain
Cloud Infrastructure (GCP/AWS), DevOps, SRE
Deliverable
production ML models | infrastructure
Required skills
Infrastructure-as-Code (IaC), Terraform, Kubernetes, Helm, Linux administration, Cloud-native architecture (GCP/AWS), Observability (Splunk, Grafana), Python, Go, Networking
Preferred skills
Terragrunt, Ansible
Technologies
Terraform, Kubernetes, Helm, ArgoCD, Splunk, Grafana, Python, Go, GCP, AWS
Responsibilities
Design and maintain scalable infrastructure using IaC; Manage and optimize Kubernetes clusters; Administer and troubleshoot Linux-based systems; Architect and manage cloud-native solutions on GCP and AWS; Implement observability practices (monitoring, logging, alerting); Develop automation tools in Python and Go; Provide production support and handle incidents; Troubleshoot complex networking issues
Seniority
Senior, hands-on IC