Senior Director of Production Engineering-Observability & Telemetry Platforms
Core
Define and execute the multi-year vision for the global Observability Platform, modernizing monitoring stacks and managing petabyte-scale telemetry pipelines.
Role type
Senior Director of Production Engineering (Observability & Telemetry)
Builds
Enterprise-wide Observability Platform and distributed Telemetry Pipelines
Domain
Cloud Security / SASE / Observability
Deliverable
production ML models | infrastructure
Required skills
Generative AI tools, AI/ML models, intelligent agents, OpenTelemetry, Kafka, Flink, Vector, VictoriaMetrics, Grafana, Prometheus, ClickHouse, OpenSearch, Kubernetes, microservices, Site Reliability Engineering, incident command structures
Preferred skills
Compliance, data retention, PII masking, data governance, chaos engineering, automated disaster recovery testing
Responsibilities
Define long-term roadmap and architectural strategy for the global Observability Platform; Architect, scale, and optimize high-throughput, low-latency telemetry ingestion and processing pipelines; Optimize platform performance and resource utilization; Partner with SRE and incident management teams to establish SLOs and enforce error budgets; Lead, mentor, and inspire a world-class engineering organization across multiple global sites; Act as a trusted partner to Product Management, Information Security, and Customer Support leadership
Seniority
Senior Director, hands-on IC with people leadership