Site Reliability Engineer - Intermediate
Core
Build and run large-scale, distributed, fault-tolerant systems ensuring reliability and performance for internal and external services.
Role type
Intermediate Site Reliability Engineer (SRE)
Builds
Production systems, containerized microservices, and automated operational workflows.
Domain
Cloud Infrastructure & DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Infrastructure as Code (Terraform), Kubernetes (GKE/EKS), public cloud management (GCP/AWS), CI/CD pipeline automation, observability (Datadog), incident triage, postmortem analysis, scripting (Groovy)
Preferred skills
Certified Kubernetes Application Developer (CKAD), GCP Associate Cloud Engineer, DevSecOps practices
Technologies
Terraform, Kubernetes, GCP, AWS, Datadog, Jenkins, GitLab CI, Docker, GitHub Actions
Responsibilities
Maintain and execute Infrastructure as Code (IaC) using Terraform across public cloud environments; Deploy, support, and troubleshoot containerized microservices operating on Kubernetes; Set up and maintain observability dashboards, alerts, and metrics; Participate in a 24/7 follow-the-sun operational rotation to manage, triage, and resolve production incidents; Collaborate with development teams to analyze system outages and execute preventative action items via postmortems.
Seniority
Intermediate, hands-on IC