Senior Storage Production Engineer - DGX Cloud
Core
Design, implement, and support large-scale storage clusters for internal and external GPU cloud services, ensuring high availability, scalability, and low-latency access for AI/ML and HPC workloads.
Role type
Senior Storage Production Engineer (Cloud Infrastructure)
Builds
Large-scale distributed storage clusters, monitoring systems, and automated storage operations for NVIDIA's GPU cloud platform.
Domain
Cloud Infrastructure / High-Performance Storage / AI/ML Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed storage systems, block/file/object storage, storage networking protocols, Linux system administration, storage automation, observability tools, infrastructure as code, incident response, capacity planning
Preferred skills
Erasure coding, replication strategies, Kubernetes/OpenStack, disaster recovery strategies, CI/CD pipelines, advanced debugging
Technologies
C/C++, Java, Python, Go, NodeJS, Bash, Ansible, Chef, Puppet, Terraform, InfluxDB, Prometheus, Grafana, Elastic stack, NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, NVMe
Responsibilities
Design and support large-scale storage clusters; develop monitoring and alerting systems; optimize storage efficiency via compression and tiering; maintain production infrastructure health; practice blameless root cause analysis; manage on-call rotation for storage systems.
Seniority
Senior, hands-on IC