Senior Software Engineer, Resilience Engineering - DGX Cloud
Core
Building org-wide reliability strategy, SLO programs, and resilience testing for NVIDIA's DGX Cloud large-scale production systems.
Role type
Senior Software Engineer - Resilience Engineering
Builds
DGX Cloud operational reliability infrastructure and tooling
Domain
Cloud computing / AI infrastructure / High-performance computing
Deliverable
production ML models | infrastructure
Required skills
Large-scale system operations, SLO program management, chaos engineering, failure injection, Go, Python, observability tooling
Preferred skills
Google SRE or Meta production engineering experience, GPU/HPC/AI training infrastructure expertise, measurable reliability improvements
Technologies
Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
Responsibilities
Build org-wide reliability strategy, stand up SLO program, lead incident response, build production code, implement chaos engineering, improve team standards
Seniority
Senior, hands-on IC