Staff Software Engineer, Observability
Core
Lead the design, deployment, and optimization of global logging, tracing, and metrics platforms for CoreWeave's AI hyperscale infrastructure.
Role type
Staff Software Engineer (Observability)
Builds
Scalable observability infrastructure (logging, tracing, metrics) for a global datacenter footprint
Domain
Cloud Infrastructure / AI Hyperscaling
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, containerization, microservices, incident management, post-mortem analysis, ClickHouse, Elastic, Loki, Victoria Metrics, Prometheus, Thanos, Grafana, data-streaming systems
Preferred skills
Running observability tools as a cloud provider, administering large-scale Kubernetes clusters
Technologies
ClickHouse, Elastic, Loki, Victoria Metrics, Prometheus, Thanos, Grafana, Kubernetes
Responsibilities
Scale logging, tracing, and metrics platforms; Develop and refine monitoring and alerting; Advise engineers on optimal usage of Observability systems; Automate interactions with Compute Infrastructure layer; Manage production clusters and deployment best practices
Seniority
Staff, hands-on IC with mentorship