CareerPlanSign in

Senior Software Engineer, Observability

New York, NY💼 Full-time🗓 2026-07-01 → 2026-09-26

Core

Design, build, and maintain core observability infrastructure (metrics, logging, tracing, telemetry) for high-performance AI workloads on GPU clusters.

Role type

Senior Software Engineer (Observability Infrastructure)

Builds

Scalable, reliable observability systems and pipelines for AI infrastructure

Domain

Cloud Infrastructure / AI / GPU Computing

Deliverable

production ML models | infrastructure

Required skills

Go, Python, Kubernetes, distributed systems, microservices, Helm, YAML, infrastructure-as-code, on-call management

Preferred skills

Observability platforms (Loki, ClickHouse, Elasticsearch, Prometheus, VictoriaMetrics, Grafana, Thanos), data streaming (Kafka), Terraform, OpenTelemetry, GPU-based AI workloads

Technologies

Go, Python, Kubernetes, Helm, YAML, Kafka, Terraform, OpenTelemetry, Loki, ClickHouse, Elasticsearch, Prometheus, VictoriaMetrics, Grafana, Thanos

Responsibilities

Design and build metrics, logging, tracing, and telemetry pipelines; develop highly reliable and scalable systems; collaborate to embed observability best practices; tackle performance and reliability challenges across thousands of GPUs; participate in on-call rotations for critical production systems

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.