Engineering Manager, Observability
Core
Lead the Observability Engineering organization to build, scale, and operate platforms for metrics, logs, traces, and telemetry pipelines that enable engineers to understand, operate, and improve production systems at scale.
Role type
Engineering Manager, Observability
Builds
Observability platforms (metrics, logs, traces, telemetry pipelines) and reliability practices for CoreWeave's AI infrastructure
Domain
Cloud Infrastructure / AI / Observability
Deliverable
production ML models | infrastructure
Required skills
Engineering management, distributed systems, reliability engineering (SLOs/SLIs), telemetry scaling, performance engineering, team hiring
Preferred skills
OpenTelemetry, Grafana, Prometheus, Kubernetes, capacity planning, high-growth environment scaling
Technologies
OpenTelemetry, Grafana, Prometheus, Kubernetes
Responsibilities
Define strategy and roadmap for observability platforms, drive platform reliability and performance improvements, guide architectural decisions, partner with infrastructure/security/application teams, hire and manage engineering teams
Seniority
Manager, hands-on technical leadership