Senior Storage Software Engineer
Core
Design and maintain high-performance parallel and distributed file systems and object storage to support massive GPU clusters for AI training and inference.
Role type
Senior Storage Software Engineer (Individual Contributor & Technical Lead)
Builds
Open-source parallel/distributed file systems, distributed object storage, and storage infrastructure for GPU fleets.
Domain
AI Infrastructure / High-Performance Computing / Storage Systems
Deliverable
production ML models | infrastructure
Required skills
C, C++, Rust, Go, Python, Linux kernel storage stacks, NVMe-oF, RDMA/RoCE/InfiniBand, I/O and metadata performance tuning, data corruption recovery, scale testing, open-source contribution
Preferred skills
Kubernetes CSI driver development, SPDK/libfabric/FUSE optimization, HPC cluster operations, AI training/inference storage experience
Technologies
Linux kernel, NVMe, SPDK, libfabric, FUSE, S3, Swift, NFS, RDMA, RoCE, InfiniBand
Responsibilities
Contribute code to open-source parallel and distributed file systems; Triage and root-cause large-scale storage issues across tens of thousands of GPUs; Validate storage architecture, performance, and durability; Define configuration and tuning standards for high-performance file systems; Partner with SRE, networking, and cloud teams on common storage architecture.
Seniority
Senior, hands-on IC with technical lead responsibilities