CareerPlanSign in

Senior Storage Production Engineer - DGX Cloud

US, CA, Santa Clara💼 Full-time💰 $176,000–$176,000🗓 2026-07-17 → 2026-09-25

Core

Design, implement, and support large-scale storage clusters for internal and external GPU cloud services, ensuring high availability, scalability, and low-latency access for AI/ML and HPC workloads.

Role type

Senior Storage Production Engineer (Cloud Infrastructure)

Builds

Large-scale distributed storage clusters, monitoring systems, and automated storage operations for NVIDIA's GPU cloud platform.

Domain

Cloud Infrastructure / High-Performance Storage / AI/ML Systems

Deliverable

production ML models | infrastructure

Required skills

Distributed storage systems, block/file/object storage, storage networking protocols, Linux system administration, storage automation, observability tools, infrastructure as code, incident response, capacity planning

Preferred skills

Erasure coding, replication strategies, Kubernetes/OpenStack, disaster recovery strategies, CI/CD pipelines, advanced debugging

Technologies

C/C++, Java, Python, Go, NodeJS, Bash, Ansible, Chef, Puppet, Terraform, InfluxDB, Prometheus, Grafana, Elastic stack, NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, NVMe

Responsibilities

Design and support large-scale storage clusters; develop monitoring and alerting systems; optimize storage efficiency via compression and tiering; maintain production infrastructure health; practice blameless root cause analysis; manage on-call rotation for storage systems.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.