Principal Observability Platform Engineer
Core
Design, build, and operate a globally consistent, regionally federated observability platform for metrics, logs, traces, events, alerts, and diagnostics to support production systems and trading teams.
Role type
Principal Observability Platform Engineer
Builds
High-scale telemetry pipelines, instrumentation libraries, dashboards, automation, and reusable platform patterns.
Domain
Financial technology / Distributed systems / Observability
Deliverable
production ML models | product features | infrastructure
Required skills
SRE, platform engineering, distributed systems, telemetry pipelines, time-series data, streaming systems, service reliability, failure mode analysis, technical trade-offs
Preferred skills
Observability tooling, developer tooling, production diagnostics, large-scale distributed systems environments
Technologies
Kafka, Grafana, ELK/OpenSearch, ClickHouse, VictoriaMetrics, InfluxDB, Telegraf, Vector, OpenTelemetry, Prometheus
Responsibilities
Design and operate components for telemetry collection, ingestion, storage, query, visualization, and alerting; Build software, APIs, and integrations to improve observability adoption; Improve scalability, reliability, and cost-effectiveness of high-volume telemetry systems; Enhance developer and operator experience through self-service workflows and tooling; Collaborate with engineering and trading teams to address production debugging needs; Own reliability and operational quality of built components; Raise standards for telemetry quality and diagnostic workflows
Seniority
Principal, hands-on IC with strategic platform ownership