Senior HPC Storage Engineer
Core
Design and implement next-gen distributed storage solutions for high-performance computing (HPC) and AI workloads, optimizing performance and cost for large-scale infrastructure.
Role type
Senior HPC Storage Engineer (Infrastructure Strategy)
Builds
Scalable distributed file, block, and object storage services for HPC clusters and AI workflows
Domain
High-Performance Computing (HPC), AI Infrastructure, Cloud Storage
Deliverable
production ML models | infrastructure
Required skills
Distributed storage architecture, Linux storage kernel development, Performance tuning, Capacity modeling, Automation tooling, Root cause analysis
Preferred skills
Parallel/distributed filesystems (Ceph, Lustre, GPFS), GPU infrastructure (CUDA, NCCL), Software Defined Networking (SDN), Deep Learning frameworks (PyTorch, TensorFlow)
Technologies
Ceph, Weka.io, Vast, Lustre, GPFS, Docker, Enroot, Python, Bash, CentOS/RHEL, Ubuntu, NVMe, HDD, SSD
Responsibilities
Research and analyze internal distributed storage services; Design and implement scalable storage services for HPC workloads; Develop tooling to automate infrastructure management and monitoring; Perform technology evaluations for distributed file systems; Collaborate with teams to capture infrastructure requirements; Support researchers with cluster performance analysis and optimization; Conduct root cause analysis for infrastructure issues
Seniority
Senior, hands-on IC