Software Engineer - Observability
Core
Design and build scalable telemetry ingest and storage pipelines, observability platforms, and diagnostic tools for multi-cloud AI infrastructure.
Role type
Senior IC software engineer (observability infrastructure)
Builds
High-throughput metrics, logs, and traces pipelines; agentic diagnostic tools; alerting and SLO infrastructure
Domain
Cloud infrastructure + AI/ML observability
Deliverable
production ML models | infrastructure
Required skills
Deep expertise in at least one observability signal (metrics, logging, tracing, error analytics); High-throughput data pipeline design; Columnar storage engine knowledge; Proficiency in Python, Rust, or Go; Experience with Prometheus, Grafana, ClickHouse, OpenTelemetry
Preferred skills
Applying AI/LLMs to operational workflows (root cause analysis, anomaly detection, intelligent alerting); Building foundational infrastructure independently
Technologies
Prometheus, Grafana, ClickHouse, OpenTelemetry, Python, Rust, Go
Responsibilities
Design scalable telemetry ingest and storage pipelines; Own and evolve core observability platforms; Build instrumentation libraries, SDKs, and integrations; Drive alerting and SLO infrastructure; Partner with Inference, Product, and Infrastructure teams
Seniority
Senior, hands-on IC