CareerPlanGet AI match score →

Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)

Seoul💼 Full-time🗓 2026-06-16 → 2026-07-31

Core

Design and evolve a planet-scale observability platform for GPU-as-a-Service infrastructure, enabling deep insights into GPU clusters, datacenter systems, and distributed ML workloads.

Role type

Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)

Builds

Production telemetry pipelines (metrics/logs/traces) and dashboards for GPUaaS infrastructure

Domain

Cloud Infrastructure / GPU Computing / Observability

Deliverable

production ML models | infrastructure

Required skills

Go, Python, Kubernetes, Linux internals, NVIDIA DCGM, CUDA ecosystem, high-performance networking (RDMA, InfiniBand), time-series databases, log storage systems, SRE principles, CI/CD automation, Terraform, OpenTelemetry

Preferred skills

Experience with GPU hardware (MIG, NVLink, PCIe), predictive observability, Zero Trust patterns, multi-tenant isolation

Technologies

Grafana Alloy, Mimir, Loki, Vector, Prometheus, Datadog, Terraform, Kubernetes, OpenTelemetry

Responsibilities

Architect low-latency, high-throughput telemetry pipelines for GPU metrics and logs; Define GPU-specific SLIs/SLOs and SLO-driven observability strategies; Build rich Grafana dashboards for fleet health and capacity planning; Lead incident forensics and cross-layer debugging for GPU contention and performance issues; Mentor engineers and drive adoption of Observability-by-Design across teams

Seniority

Sr. Staff, strategic architecture & hands-on engineering

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗