Senior AI infrastructure engineer - EDA Infrastructure
Core
Build automated, intelligence-driven systems for telemetry, operational excellence, and hardware/software cataloging to protect NVIDIA's critical AI platforms and GPU cloud services.
Role type
Senior IC infrastructure engineer (observability & inventory)
Builds
Scalable telemetry pipelines, automated incident workflows, and trusted hardware/software catalogs for GPU cloud services
Domain
High-performance computing + AI infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, TypeScript, Java, software engineering principles, production environment experience, cross-functional leadership
Preferred skills
incident-management processes, ML model deployment, AI agent frameworks, observability platforms, CMDB
Technologies
Python, Go, TypeScript, Java
Responsibilities
Build and operate scalable telemetry pipelines for metrics, logs, and traces; Standardize and automate incident and maintenance workflows; Build and maintain physical hardware and software catalogs
Seniority
Senior, hands-on IC