Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Core
Architecting and operating multi-petabyte distributed storage systems and Kubernetes-native platforms to support extreme-scale AI training and inference workloads.
Role type
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Builds
Multi-petabyte AI/ML storage systems, Kubernetes storage operators, and self-service provisioning platforms for GPU clusters.
Domain
Artificial Intelligence Infrastructure / High-Performance Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed storage systems (Ceph, WekaFS, Lustre, Vast), Kubernetes storage operators, Go, Python, RDMA/InfiniBand networking, Linux storage stack (ext4, xfs, LVM, NVMe), Terraform, Ansible, Helm, GitOps, Prometheus, Grafana, fio, iperf3
Preferred skills
GPU Direct Storage (GDS), NVMe-oF, ML/AI storage patterns (checkpointing, dataset caching), advanced storage benchmarking and profiling
Technologies
Vast, Weka, Ceph, Lustre, Kubernetes, Go, Python, Terraform, Ansible, Helm, ArgoCD, Prometheus, Grafana, Thanos, fio, iperf3
Responsibilities
Architect and implement the technical strategy and storage roadmap for scaling GPU fleets; Engineer and scale multi-petabyte AI/ML storage systems with automated tiering; Develop intelligent caching and tiered storage architectures for extreme IOPS; Tune storage isolation at L2/L3 network layers for multi-tenancy; Code Kubernetes storage operators for automated provisioning and quota enforcement; Engineer end-to-end data paths achieving 10+ GB/s per GPU node; Optimize data paths through benchmarking and contribute to open-source storage projects.
Seniority
Staff, hands-on IC with technical leadership