Site Reliability Engineer (SRE) (GovTech)
Core
Design and operate GitLab, AWS, and Kubernetes-based infrastructure to ensure the stability, scalability, and performance of a Whole-of-Government runtime platform.
Role type
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
Builds
GitLab, AWS, and Kubernetes-based runtime platform for government tenants
Domain
Government Technology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes orchestration, AWS cloud services, Infrastructure as Code (Terraform), CI/CD pipeline automation, Observability (Prometheus, Grafana, ELK), Incident management, Go/Python/Bash scripting, Containerization (Helm, Kustomize), Networking and Security in cloud contexts
Preferred skills
Kubernetes Operators, Service Mesh, Chaos Engineering, CKA/CKAD certifications
Technologies
GitLab, AWS, Kubernetes, EKS, Terraform, Prometheus, Grafana, ELK Stack, Helm, Kustomize, Go, Python, Bash, ArgoCD
Responsibilities
Develop automation via CI/CD pipelines to reduce toil; Implement observability solutions and self-remediation; Participate in on-call rotations and conduct post-incident reviews; Design secure and compliant solutions; Resolve performance bottlenecks and track KPIs; Act as a technical advisor for tenants on containerization and cloud-native best practices; Develop and maintain playbooks and documentation.
Seniority
Senior, hands-on IC