Staff Engineer, Storage Engine
Core
Designing and implementing distributed storage solutions for exabyte-scale, S3-compatible object storage to support data-intensive AI workloads.
Role type
Staff Engineer, Storage Engine
Builds
Reliable, scalable storage solutions and dedicated storage clusters for AI labs, startups, and global enterprises.
Domain
Cloud infrastructure, AI compute, distributed systems
Deliverable
production ML models | infrastructure
Required skills
Distributed storage systems, Object storage (S3), Distributed filesystems (Ceph, DAOS), Systems programming (Go, C, Rust), Storage protocols (NFS, FUSE), Cloud-native infrastructure (Kubernetes), Telemetry pipelines (ClickHouse, Prometheus, Grafana)
Preferred skills
AI tools for software development, Cross-functional collaboration, Mentorship
Technologies
RDMA, GPU Direct Storage, Ceph, DAOS, S3, NFS, FUSE, Kubernetes, ClickHouse, Prometheus, Grafana
Responsibilities
Design and implement distributed storage solutions; Integrate dedicated storage clusters into customer environments; Optimize storage performance using RDMA and GPU Direct Storage; Improve reliability, durability, security, and observability of the storage stack; Monitor and troubleshoot production storage systems; Develop metrics and dashboards for storage visibility; Analyze telemetry to drive improvements in throughput and latency; Mentor engineers on building high-performance distributed systems.