Senior Storage Production Engineer - DGX Cloud
Core
Design, implement, and support large-scale storage clusters for internal and external-facing GPU cloud services, ensuring high availability, scalability, and low-latency data access for HPC and AI/ML workloads.
Role type
Senior Storage Production Engineer
Builds
Large-scale distributed storage systems and GPU cloud services
Domain
High-Performance Computing (HPC) and Artificial Intelligence (AI/ML)
Deliverable
production ML models | infrastructure
Required skills
Distributed storage architecture, high-performance file/object storage, storage networking protocols (NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, NVMe), Linux system administration, automation scripting (C/C++, Java, Python, Go, Bash), infrastructure as code (Ansible, Chef, Puppet, Terraform), observability (InfluxDB, Prometheus, Grafana, Elastic stack), capacity planning, performance tuning, erasure coding, disaster recovery
Preferred skills
Kubernetes/OpenStack/hybrid cloud architectures, Git/CI/CD pipelines, systematic debugging, automated fault detection and remediation
Technologies
Kubernetes, OpenStack, Ansible, Chef, Puppet, Terraform, InfluxDB, Prometheus, Grafana, Elastic stack, C/C++, Java, Python, Go, NodeJS, Bash, NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, NVMe
Responsibilities
Design and support large-scale storage clusters; develop monitoring and alerting systems; optimize storage efficiency via compression and tiering; maintain production infrastructure health; practice incident response and root cause analysis; scale systems using AI-driven automation.
Seniority
Senior, hands-on IC