CareerPlanGet AI match score →

Infrastructure Engineer (Observability)

New York, New York🌐 Remote💼 Full-time💰 $180,000–$180,000🗓 2026-05-27 → 2026-07-31

Core

Build and evolve scalable observability platforms (metrics, logs, traces) for large-scale GPU-enabled bare-metal infrastructure, serving internal teams and external customers.

Role type

Senior Infrastructure Engineer (Observability)

Builds

Multi-tenant observability systems, telemetry pipelines, and productized monitoring experiences for GPU/HPC clusters.

Domain

Cloud Infrastructure / AI Systems / Observability

Deliverable

production ML models | infrastructure

Required skills

Observability platform design, telemetry pipeline engineering, alerting system design, Python/Go/Bash, Kubernetes observability, streaming data processing, multi-tenant architecture

Preferred skills

GPU observability (NVIDIA DCGM), InfiniBand fabric monitoring, correlation engines, predictive alerting, large-scale HPC cluster management

Technologies

Prometheus, Grafana, ELK, VictoriaMetrics, Kafka, OTEL, Promtail, Kubernetes, NVIDIA DCGM, InfiniBand

Responsibilities

Design and operate telemetry pipelines ingesting data from GPUs, CPUs, and networking; implement noise-resistant alerting systems; create dashboards for InfraOps and Customer Success; partner with platform teams to embed observability into core systems.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗