CareerPlanSign in

Software Engineer (Site Reliability)

Sydney, New South Wales💼 Full-time🗓 2026-09-17 → 2026-09-29

Core

First dedicated Site Reliability Engineer owning reliability and core platform decisions for scaling an AI-powered interactive entertainment platform to hundreds of millions of users.

Role type

Senior Site Reliability Engineer (Platform)

Builds

High-scale GPU clusters, AWS infrastructure, and observability systems for AI generation services.

Domain

AI/ML Infrastructure, Cloud Computing, Interactive Entertainment

Deliverable

production ML models | infrastructure

Required skills

AWS infrastructure-as-code, Kubernetes/ECS, observability (metrics/tracing/alerting), incident response, CI/CD pipeline design, GPU cluster orchestration, SLO enforcement, cloud cost optimization

Preferred skills

TypeScript, Next.js, React, Postgres, Temporal

Technologies

AWS, Kubernetes, ECS, GPU clusters, Postgres, Temporal, tRPC, Next.js, React, TailwindCSS

Responsibilities

Improve uptime and reduce RTO for critical services, orchestrate and harden GPU clusters, implement platform-wide observability, optimize AWS infrastructure and reduce cloud spend

Seniority

Senior, hands-on IC

Sourced via viewjobs · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.