Infrastructure Engineer (Observability)
Core
Build and evolve scalable observability platforms (metrics, logs, traces) for large-scale GPU-enabled bare-metal infrastructure, serving internal teams and external customers.
Role type
Senior Infrastructure Engineer (Observability)
Builds
Multi-tenant observability systems, telemetry pipelines, and productized monitoring experiences for GPU/HPC clusters.
Domain
Cloud Infrastructure / AI Systems / Observability
Deliverable
production ML models | infrastructure
Required skills
Observability platform design, telemetry pipeline engineering, alerting system design, Python/Go/Bash, Kubernetes observability, streaming data processing, multi-tenant architecture
Preferred skills
GPU observability (NVIDIA DCGM), InfiniBand fabric monitoring, correlation engines, predictive alerting, large-scale HPC cluster management
Technologies
Prometheus, Grafana, ELK, VictoriaMetrics, Kafka, OTEL, Promtail, Kubernetes, NVIDIA DCGM, InfiniBand
Responsibilities
Design and operate telemetry pipelines ingesting data from GPUs, CPUs, and networking; implement noise-resistant alerting systems; create dashboards for InfraOps and Customer Success; partner with platform teams to embed observability into core systems.
Seniority
Senior, hands-on IC