Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
Core
Keep GitLab's user-facing services and production systems reliable, scalable, and efficient by building automation, operating Kubernetes clusters, and managing incident response.
Role type
Senior IC Site Reliability Engineer (Infrastructure Platforms)
Builds
Production infrastructure tooling, automation workflows, and observability stacks for GitLab.com and dedicated services.
Domain
SaaS / DevSecOps / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes operations, Infrastructure as Code (Terraform), Go/Ruby code debugging, Cloud provider expertise (AWS/GCP), Observability (metrics/logs/SLOs), Incident response, Automation development
Preferred skills
AI integration for toil reduction, Async distributed work, Manager-of-one leadership
Technologies
Kubernetes, Terraform, Go, Ruby, AWS, GCP
Responsibilities
Operate and troubleshoot production systems on Kubernetes, Write and maintain infrastructure as code, Participate in on-call and incident response, Build automation to reduce manual toil, Contribute to the observability stack, Document architecture decisions and runbooks
Seniority
Senior, hands-on IC