Principal Distributed Systems Engineer - Observability
Core
Architect and build a multi-petabyte scale distributed tracing platform and big-data pipeline to power observability, anomaly detection, and AI-driven incident triage.
Role type
Principal Distributed Systems Engineer (Observability)
Builds
Distributed tracing infrastructure, Kafka-based ingestion pipelines, Iceberg-on-S3 storage, and AI-assisted incident management features.
Domain
Cloud-native observability, large-scale data processing, distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems architecture, Big data pipeline design, High availability/Disaster recovery, System security, Algorithmic thinking, Technical leadership
Preferred skills
Observability AI integration, Cloud-native technology evaluation
Technologies
ClickHouse, Grafana Tempo, Kafka, Spark, Flink, Iceberg, S3, Elasticsearch, AWS
Responsibilities
Architect multi-petabyte tracing platform, Own big-data pipeline (Kafka/Spark/Flink/Iceberg), Drive performance and scaling, Lead HA/DR design, Design security architecture, Own operational excellence, Evaluate new technologies, Shape Observability AI strategy, Mentor engineers
Seniority
Principal, hands-on IC with strategic direction