Site Reliability Engineer
Core
Deliver and sustain mission-critical technology solutions for Defence, Intelligence, and Government sectors, ensuring reliability, performance, and resilience of complex production environments.
Role type
Site Reliability Engineer (SRE)
Builds
Mission-critical software solutions and infrastructure for national security outcomes
Domain
Defence, Intelligence, and Government sectors
Deliverable
production ML models | infrastructure
Required skills
Linux/UNIX administration, Python/Go/Bash scripting, Infrastructure as Code (Terraform, Ansible, Puppet, Chef), monitoring and observability, CI/CD pipeline management, cloud platform support (VMware, AWS, Azure, GCP), incident response and root cause analysis
Preferred skills
Experience in Defence or National Security environments, fault-tolerant large-scale system architecture, secure systems administration, DevOps practices
Technologies
Terraform, Ansible, Puppet, Chef, Python, Go, Bash, Jenkins, GitLab CI, VMware, AWS, Azure, GCP
Responsibilities
Support availability, capacity, and reliability of complex production environments; Provide Level 2 and Level 3 operational support and lead incident response; Develop and maintain monitoring, alerting, and logging capabilities; Enhance automation, infrastructure-as-code, and CI/CD practices
Seniority
Mid-Senior, hands-on IC