CareerPlanSign in

Reliability & Observability Analyst II

Sydney, New South Wales💼 Full-time🗓 2026-09-07 → 2026-09-26

Core

Support 24/7 HPC Data Center Operations by performing advanced incident triage, improving alert quality, and maintaining operational telemetry for GPU clusters.

Role type

Senior IC Reliability & Observability Analyst (HPC Data Center)

Builds

Actionable operational telemetry, tuned alerting systems, and incident response workflows for AI training/inference clusters

Domain

AI Cloud Infrastructure / High-Performance Computing / Data Center Operations

Deliverable

production ML models | product features | dashboards & analysis | client delivery

Required skills

incident lifecycle management, alert tuning and routing, Linux systems administration, GPU health analysis, log/metrics/alert correlation, AIOps validation, small automation scripting, SLI/SLO reporting, ITSM tooling, RCA workflows

Preferred skills

experience with high-density GPU clusters, mentoring analysts, cross-functional partnership with engineering

Technologies

ServiceNow, Jira, Splunk, Datadog, Linux, GPU clusters

Responsibilities

Perform advanced incident triage and escalation; tune monitoring dashboards and alert quality; correlate logs, metrics, and alerts across distributed systems; validate AIOps outputs during incidents; write small automation artifacts to reduce toil; maintain operational reports and incident records

Seniority

Mid-Senior, hands-on IC

Sourced via viewjobs · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.