Staff Site Reliability Engineer, Environment Automation
Core
Operate and automate hundreds of isolated GitLab tenant environments, ensuring they remain secure, consistent, and reliable at scale.
Role type
Staff Site Reliability Engineer (Environment Automation)
Builds
Multi-tenant GitLab Dedicated instances, infrastructure-as-code workflows, and observability systems
Domain
DevSecOps platform infrastructure, cloud-native environments (GCP, AWS)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Terraform mastery, Kubernetes production operations, Go/Ruby code analysis, incident response leadership, multi-tenant architecture design, observability stack management
Preferred skills
Ansible, Jsonnet templating, Helm Charts, cloud provider IAM/networking/storage integration
Technologies
Terraform, Ansible, Kubernetes, Go, Ruby, Prometheus, ELK, Grafana, Helm, GCP, AWS
Responsibilities
Design and implement automation for provisioning and managing hundreds of isolated environments; troubleshoot production issues across clusters and cloud services; replace manual workflows with infrastructure-as-code solutions; build observability systems to detect bottlenecks and predict capacity; lead incident response and postmortem efforts; influence architectural decisions around automation and scalability
Seniority
Staff, hands-on IC with strategic influence