Senior Manager, Observability
Core
Lead the Observability Engineering organization to build, scale, and operate platforms for metrics, logs, traces, and telemetry pipelines that enable engineers to understand, operate, and improve production systems at scale.
Role type
Senior Manager, Observability Engineering
Builds
Observability platforms (metrics, logs, traces, telemetry pipelines) for CoreWeave's AI infrastructure
Domain
Cloud Infrastructure / AI / Observability
Deliverable
production ML models | infrastructure
Required skills
Software engineering at scale, Engineering management, Observability platform design (logs, metrics, traces, alerting), Reliability engineering (SLOs, SLIs, incident management, error budgets), Telemetry system scaling, Distributed systems architecture, Team hiring and management
Preferred skills
OpenTelemetry, Grafana, Prometheus-compatible systems, Kubernetes, Cloud-native infrastructure, AI/ML infrastructure support, Capacity planning for high-ingest systems
Technologies
OpenTelemetry, Grafana, Prometheus, Kubernetes
Responsibilities
Define strategy and roadmap for observability platforms, Drive platform reliability and performance improvements, Guide architectural decisions across observability infrastructure, Partner with infrastructure, security, and application teams to improve instrumentation, Manage and grow the engineering team
Seniority
Senior, hands-on IC with management responsibilities