CareerPlanSign in

Senior Site Reliability Engineer

Australia💼 Full-time🗓 2026-06-04 → 2026-09-25

Core

Operating and maintaining production clusters, developing observability solutions, and ensuring platform reliability through automation and monitoring.

Role type

Senior Site Reliability Engineer (Cloud Platform)

Builds

Production clusters, observability solutions (metrics, logs, traces), automation strategies, and runbooks.

Domain

Cloud infrastructure, container orchestration, observability

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Kubernetes or Docker Swarm cluster operation, AWS, ELK stack, Prometheus/Grafana, Bash/Python/Go scripting, incident response and root cause analysis, runbook development, embedding reliability in CI/CD

Preferred skills

Linux administration (RedHat/CentOS/AL), Infrastructure as Code (Terraform/Ansible), TCP/IP and networking concepts, RDBMS (MySQL/Postgres), NoSQL (Redis), telephony knowledge (SIP/VoIP)

Technologies

Kubernetes, Docker Swarm, AWS, ELK, Prometheus, Grafana, Bash, Python, Go, Terraform, Ansible, MySQL, Postgres, Redis

Responsibilities

Ensure platform reliability and availability via proactive monitoring and automation; serve as first responder for incidents and lead root cause analysis; develop troubleshooting documentation and optimized runbooks; collaborate with engineering teams to embed reliability into the software delivery lifecycle; design and evolve observability solutions; participate in on-call rotations and improve alert quality; champion a culture of reliability and continuous improvement.

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.