AI Infrastructure Engineer (Storage)
Core
Design, deploy, and maintain high-performance distributed and tiered storage systems for large-scale AI and data pipelines.
Role type
Senior Infrastructure Engineer (Storage)
Builds
Production storage platforms supporting AI/ML workloads, container images, and hybrid cloud environments.
Domain
AI Infrastructure / Distributed Storage Systems
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Ceph cluster management, distributed filesystems (Lustre, BeeGFS), tiered storage architecture, Bash/Python scripting, container image management, cloud integration (AWS/Azure/GCP), data security best practices
Preferred skills
GPU compute environment knowledge, observability tools (Prometheus, Grafana), open-source storage contributions, object storage systems (S3, MinIO)
Responsibilities
Design and implement storage platforms for AI pipelines, manage distributed storage systems, oversee tiered storage architectures, ensure data integrity and security, develop automation tools, integrate with public cloud services, troubleshoot storage issues, contribute to architecture improvements
Seniority
Senior, hands-on IC