CareerPlanGet AI match score →

Senior AI and HPC Observability Engineer

2 Locations💼 Full-time💰 $152,000–$152,000🗓 2026-03-02 → 2026-07-31

Core

Design and scale next-generation observability and telemetry platforms for high-volume metrics, logs, and traces across distributed AI/HPC environments.

Role type

Senior IC backend/distributed systems engineer (observability)

Builds

High-throughput telemetry pipelines, backend services, and monitoring frameworks for AI superclusters

Domain

High-performance computing (HPC) and Artificial Intelligence (AI)

Deliverable

production ML models | infrastructure

Required skills

Python, Go, Java, distributed systems design, time-series data systems, streaming technologies, Kubernetes, fault-tolerant system design

Preferred skills

OpenTelemetry, Prometheus, Kafka, Spark, Flink, GPU workload monitoring, statistical anomaly detection

Technologies

OpenTelemetry, Prometheus, Kafka, Spark, Flink, Kubernetes, PromQL

Responsibilities

Design and scale observability platforms handling high-volume metrics, logs, and traces; Build high-performance backend services for telemetry ingestion, processing, and routing; Develop and extend OpenTelemetry collectors, processors, exporters, and instrumentation libraries; Build and optimize metrics pipelines using large-scale time-series storage systems; Design and operate real-time and batch telemetry pipelines using streaming and distributed data technologies; Improve platform reliability, performance, and cost efficiency through tuning, capacity planning, and system optimization; Develop monitoring, alerting, and service reliability frameworks to ensure platform health and performance

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗