HPC Storage Engineer - West Coast
Core
Design, scale, and ensure the reliability of a multi-region distributed storage ecosystem (network volumes, local NVMe, S3-compatible object storage) supporting AI training, fine-tuning, and inference workloads.
Role type
Senior hands-on IC storage engineer
Builds
Distributed storage deployments, control-plane services, provisioning workflows, and data movement pipelines
Domain
AI infrastructure / Cloud storage systems
Deliverable
production ML models | infrastructure
Required skills
Distributed storage systems (Ceph, MinIO, Lustre, GPFS, ZFS), Linux internals (block layer, filesystems, NVMe), Networking for storage (RDMA/RoCE, IB/Ethernet, MTU, jumbo frames), Go/Python/Rust, Observability (Prometheus, Grafana, Datadog), Capacity planning, Automation/Infrastructure as Code
Preferred skills
AI/ML storage patterns (checkpointing, dataset streaming, GPUDirect Storage), Kubernetes storage internals (CSI drivers, PV/PVC), Bare-metal/colocation management, Multi-tenant isolation
Technologies
Ceph, MinIO, Lustre, GPFS, S3, Kubernetes, CSI, Prometheus, Grafana, Datadog, Go, Python, Rust, NVMe, RDMA, RoCE
Responsibilities
Own capacity, durability, availability, and performance of storage systems; Tune I/O paths and network fabrics; Diagnose hard performance problems; Lead capacity expansions and migrations; Write production code for control-plane and automation; Build dashboards, SLOs, and alerts; Participate in on-call rotation
Seniority
Senior, hands-on IC