Principal Platform Engineer, Observability (CIPE)
Core
Architect, build, and evolve a modern observability platform integrating AI agents and developer workflows to support metrics, logs, traces, and incident response across infrastructure and applications.
Role type
Principal Platform Engineer (Observability)
Builds
Scalable observability stack, AI-enabled workflows, instrumentation libraries, and developer tools for telemetry collection and analysis.
Domain
Cybersecurity, Cloud Infrastructure, Observability, AI Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
OpenTelemetry, Prometheus, Kubernetes, Python, distributed systems, AI agent orchestration, SRE principles, alerting frameworks, synthetic monitoring
Preferred skills
Go, Java, Rust, Node.js, Chronosphere, Grafana, Terraform, GitOps
Technologies
OpenTelemetry, Prometheus, Jaeger, Alertmanager, Grafana, Kubernetes, Helm, Python, Claude, Codex, MCP servers
Responsibilities
Design architecture standards for telemetry collection, processing, storage, and visualization; Lead adoption of OpenTelemetry and AI coding tools; Build scalable systems for metrics, tracing, and synthetic monitoring; Mentor engineers and drive adoption of observability patterns; Integrate AI agents into observability workflows for triage and remediation.
Seniority
Principal, hands-on IC with strategic leadership