Principal Site Reliability Engineer (Sovereign Cloud)
Core
Design, build, and operate reliable, secure cloud infrastructure for Sovereign Cloud services, focusing on automation, observability, and high availability.
Role type
Principal Site Reliability Engineer
Builds
Cloud infrastructure, automation frameworks, and monitoring/alerting systems for internal services
Domain
Cybersecurity, Cloud Infrastructure, Site Reliability Engineering
Deliverable
infrastructure
Required skills
Site Reliability Engineering, Cloud Infrastructure (AWS/GCP), Kubernetes, Infrastructure as Code (Terraform/Ansible), Python, Shell Scripting, Linux Administration, Distributed Systems Troubleshooting, CI/CD, Observability
Preferred skills
GitLab CI, ArgoCD, Go, Java
Technologies
Terraform, Kubernetes, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, Go, Ansible, Helm
Responsibilities
Design and operate cloud infrastructure, develop automation tools, orchestrate monitoring and alerting, lead root cause analysis, participate in on-call rotations, ensure application scalability and reliability
Seniority
Principal, hands-on IC with leadership in technical design and troubleshooting