Senior System Software Engineer - Data Platform Observability
Core
Architect and build a centralized, high-performance Data & Observability Platform for NVIDIA's internal AI, HW, and SW engineering teams to visualize chip telemetry, debug distributed pipelines, and ensure platform reliability.
Role type
Senior System Software Engineer (Data Platform & Observability)
Builds
Centralized telemetry pipelines, policy enforcement systems, web interfaces/APIs, and automated workflow orchestration for observability.
Domain
Semiconductor / AI Infrastructure / Observability
Deliverable
production ML models | product features | infrastructure
Required skills
High-performance backend systems programming, full-stack development, time-series database internals, inverted indexes, Kubernetes stateful services, event streaming, data lake formats, policy-as-code, data governance
Preferred skills
Custom Grafana plugin development, legacy monolith to microservices migration, Vector-based pipelines, OpenTelemetry (OTEL) collector configuration, instrumentation SDKs
Technologies
Python, JS, Java, Rust, Go, React, Apache Spark, Elastic/Open Search, Grafana, Prometheus, Helm, Terraform, Ansible
Responsibilities
Design centralized telemetry pipelines handling massive scale; implement data governance and access control enforcement; develop modern web interfaces and APIs; implement cost-effective tiered storage architectures; architect workflow orchestration for platform maintenance and data lifecycle management.
Seniority
Senior, hands-on IC