Software Engineer, Observability
Core
Design, build, and maintain scalable systems that process and surface telemetry data (logging, tracing, metrics) across distributed environments to enable monitoring and optimization of AI workloads on GPU-dense infrastructure.
Role type
Software Engineer, Observability
Builds
Logging, tracing, and metrics platforms; high-throughput data pipelines for GPU-dense AI infrastructure
Domain
Cloud Infrastructure / AI / Observability
Deliverable
production ML models | product features | infrastructure
Required skills
Go, Python, Kubernetes, containerization, microservices, observability systems (metrics, logging, tracing), incident response
Preferred skills
ClickHouse, Elastic, Loki, VictoriaMetrics, Prometheus, Thanos, OpenTelemetry, Grafana, Terraform, canary/blue-green deployment, Kafka, MLOps tooling
Responsibilities
Design and build scalable telemetry systems; improve system reliability through monitoring and alerting; participate in on-call rotations; optimize performance across large-scale infrastructure
Seniority
Mid-level, hands-on IC