Lead Site Reliability Engineer
Core
Lead Site Reliability Engineer responsible for ensuring 99.9%+ uptime, managing SLOs/SLIs, and driving reliability for critical services.
Role type
Lead Site Reliability Engineer
Builds
resilient, scalable, high-performance services and infrastructure
Domain
Cloud infrastructure, distributed systems, marketing technology
Deliverable
production ML models | infrastructure
Required skills
SLO/SLI management, incident response, root cause analysis, capacity planning, observability design, chaos engineering, Infrastructure as Code, Linux systems administration, cloud platform expertise, container orchestration, programming in Python or Go, shell scripting, distributed tracing integration
Preferred skills
distributed systems experience, statistical analysis of metrics, high-performance low-latency systems knowledge, on-call rotation experience, production-level code writing, Chaos Engineering drill execution
Technologies
AWS, Kubernetes, EKS, Fargate, OpenTelemetry, Honeycomb, Grafana, Prometheus, Thanos, ELK, Loki, Terraform, Pulumi, Chaos Mesh, AWS Fault Injection Simulator
Responsibilities
Implement and manage SLOs, SLIs, and error budgets; Develop resilient systems ensuring 99.9%+ uptime; Lead incident response and post-incident reviews; Automate incident detection and response; Write software to support reliability needs; Design and implement full observability; Perform capacity planning and performance testing; Collaborate on building reliable services; Ensure best practices in infrastructure design and deployment; Champion Infrastructure as Code; Participate in chaos engineering initiatives; Participate in on-call rotation; Drive advanced alerting and anomaly detection
Seniority
Senior, hands-on IC with leadership responsibilities