Staff Engineer, Network Observability
Core
Define and evolve technical direction for network observability, building resilient telemetry systems and closed-loop automation workflows to provide fast, trustworthy insight into network behavior.
Role type
Staff Engineer, Network Observability
Builds
Scalable observability platform (collectors, persistence, visualization, alerting) for CoreWeave's infrastructure
Domain
Cloud Infrastructure / Network Engineering
Deliverable
production ML models | infrastructure
Required skills
Network observability architecture, scalable telemetry design, cross-team technical leadership, systems thinking, incident response, mentorship
Preferred skills
Machine learning for anomaly detection, distributed tracing (OpenTelemetry/Jaeger/Zipkin), technical roadmap planning, network certifications
Technologies
gNMI, SNMP, Prometheus, Loki, Clickhouse, Grafana, Alertmanager, Python, Go, Bash, Ansible, Jinja2, Kubernetes, SONiC, HPE Junos, NVIDIA Cumulus Linux, Nokia SR OS, SR Linux
Responsibilities
Set technical direction for network observability across multiple teams; Lead design of scalable observability solutions using diverse collector and persistence technologies; Drive standardization of observability patterns across the network stack; Partner with leadership to prioritize investments and make high-leverage tradeoffs; Act as technical expert for critical observability challenges and incidents; Mentor junior and senior engineers; Participate in architectural decisions and RFCs; Join rotating on-call schedule as senior escalation point
Seniority
Staff, hands-on IC with strategic influence