Site Reliability / Gitops Engineer
Core
Drive operations automation and maintain IT production services for over 60 million Ubuntu users across private and public clouds.
Role type
Senior Site Reliability / GitOps Engineer
Builds
Automated infrastructure as code (IaC) practices, resilient cloud/container portfolios, and operational playbooks.
Domain
Open source operating systems, public cloud, and distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
Infrastructure as Code (IaC), Python software development, Linux networking and storage administration, CI/CD pipelines, observability tooling (Prometheus, Grafana, Elasticsearch), troubleshooting from kernel to web
Preferred skills
Experience with Ubuntu or Debian, knowledge of Ceph storage, familiarity with open-source ecosystems
Technologies
Python, Linux, Kubernetes, Terraform, Prometheus, Grafana, Elasticsearch, Ceph
Responsibilities
Develop and improve IaC processes for automation across clouds, maintain operational responsibility for core services and networks, design service architecture and operational procedures, troubleshoot complex distributed systems, mentor team members on best practices, carry final responsibility for time-critical escalations
Seniority
Senior, hands-on IC