Senior Solutions Architect, AI Factory Observability and Visualization - NVIS
Core
Develops full-spectrum visibility for HPC systems and AI factories by transforming intricate telemetry into actionable perspectives.
Role type
Senior Solutions Architect (HPC/AI Observability)
Builds
Observability solutions, telemetry pipelines, and dashboards for HPC and AI factory systems.
Domain
High-Performance Computing (HPC) and AI Infrastructure
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Linux system management, multi-GPU/multi-node cluster architecture, Python scripting, Shell/Bash automation, observability stack expertise (Prometheus, Grafana, Loki), GPU/fabric telemetry analysis, distributed systems knowledge
Preferred skills
AI factory infrastructure build/deployment, SPC systems engineering, data pipeline construction for scale
Technologies
Python, Shell, Bash, Prometheus, Grafana, Loki, DCGM, NVLink, InfiniBand, Ethernet
Responsibilities
Run AI factory validation tools and microbenchmarks to assess system health; establish system health metrics and thresholds; build and extend telemetry surfaces across hardware, fabric, and workload; develop automation for data collection and presentation; investigate visibility gaps; collaborate with cross-functional teams to ready systems for customer deployment
Seniority
Senior, hands-on IC