Senior Site Reliability Engineer
Core
Maintain reliability, performance, and stability for a high-traffic real-time production system used by millions.
Role type
Senior hands-on Site Reliability Engineer (Infrastructure)
Builds
Production infrastructure, observability tooling, and CI/CD pipelines for a high-load platform
Domain
Cloud Infrastructure / DevOps / High-availability Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (EKS), GitOps (Flux/ArgoCD), Terraform, AWS, Docker, CI/CD, Monitoring/Observability (Datadog, Prometheus, Grafana, ELK, CloudWatch), Networking, Scripting (Python/Go/Node.js), Git, Incident Management (PagerDuty/Opsgenie)
Preferred skills
None explicitly stated
Technologies
Kubernetes, EKS, Terraform, Helm, Flux, ArgoCD, AWS, Datadog, Prometheus, Grafana, ELK, CloudWatch, PagerDuty, Opsgenie
Responsibilities
Monitor platform health and manage alerts; Respond to incidents and perform root cause analysis; Build and improve monitoring and observability; Deploy and optimize infrastructure via IaC; Maintain CI/CD pipelines; Collaborate with engineering teams on deployments
Seniority
Senior, hands-on IC