CareerPlanGet AI match score →

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

San Francisco🌐 Remote💼 Full-time💰 $250,000–$250,000🗓 2026-07-10 → 2026-07-31

Core

Architecting and operating multi-petabyte distributed storage systems and Kubernetes-native platforms to support extreme-scale AI training and inference workloads.

Role type

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Builds

Multi-petabyte AI/ML storage systems, Kubernetes storage operators, and self-service provisioning platforms for GPU clusters.

Domain

Artificial Intelligence Infrastructure / High-Performance Computing / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Distributed storage systems (Ceph, WekaFS, Lustre, Vast), Kubernetes storage operators, Go, Python, RDMA/InfiniBand networking, Linux storage stack (ext4, xfs, LVM, NVMe), Terraform, Ansible, Helm, GitOps, Prometheus, Grafana, fio, iperf3

Preferred skills

GPU Direct Storage (GDS), NVMe-oF, ML/AI storage patterns (checkpointing, dataset caching), advanced storage benchmarking and profiling

Technologies

Vast, Weka, Ceph, Lustre, Kubernetes, Go, Python, Terraform, Ansible, Helm, ArgoCD, Prometheus, Grafana, Thanos, fio, iperf3

Responsibilities

Architect and implement the technical strategy and storage roadmap for scaling GPU fleets; Engineer and scale multi-petabyte AI/ML storage systems with automated tiering; Develop intelligent caching and tiered storage architectures for extreme IOPS; Tune storage isolation at L2/L3 network layers for multi-tenancy; Code Kubernetes storage operators for automated provisioning and quota enforcement; Engineer end-to-end data paths achieving 10+ GB/s per GPU node; Optimize data paths through benchmarking and contribute to open-source storage projects.

Seniority

Staff, hands-on IC with technical leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗