数据加速高级开发工程师(深圳/北京/上海/杭州)
Core
Design and develop distributed storage and caching systems optimized for AI training, inference, and compute scenarios.
Role type
Senior IC distributed systems engineer (AI infrastructure)
Builds
High-performance distributed storage and caching systems for AI workloads
Domain
AI infrastructure, distributed systems, storage, and hardware acceleration
Deliverable
production ML models
Required skills
C++, Python, Linux backend development, distributed systems, caching, storage systems, high-concurrency services, algorithms and data structures, RDMA, FUSE, file systems, GPU training/inference frameworks
Preferred skills
Open source contributions, top-tier conference publications, deep expertise in specific sub-domains (hardware acceleration, model training)
Technologies
Alluxio, 3FS, CEPH, AWS-S3, JuiceFS, RDMA, FUSE, GPU
Responsibilities
Architect and develop distributed storage/caching systems for AI scenarios; Optimize system performance, stability, and data access efficiency in large-scale production; Optimize IO paths for AI training/inference to maximize hardware utilization; Research and implement cutting-edge technologies in the field.