Senior Site Reliability Engineer
Core
Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues.
Role type
Senior Site Reliability Engineer (IC with mentorship)
Builds
Production services, alerting pipelines, dashboards, SLO-driven monitoring strategies, and internal tooling
Domain
Cloud infrastructure, observability, and AI-driven operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Docker, Linux (kernel internals, TCP/IP, DNS, load balancers), Python, Ansible, Infrastructure as Code (Terraform/Pulumi), Icinga, Prometheus, Grafana, LLMs/embeddings/ML pipelines
Preferred skills
CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions), Service Level Objectives/Indicators/Agreements management
Technologies
Kubernetes, Docker, Icinga, Prometheus, Grafana, Python, Ansible, Terraform, Pulumi, Jenkins, GitLab CI, GitHub Actions
Responsibilities
Design and operate scalable production services; Build and maintain alerting pipelines and SLO-driven monitoring; Lead incident response and root-cause analysis; Develop Infrastructure as Code and internal tooling; Mentor junior SREs; Apply LLM-driven log analysis and AI tools for incident response
Seniority
Senior, hands-on IC with mentorship
