训练推理一体化存储研发高级工程师-Data
Core
Design and implement multi-level storage systems for recommendation/advertising large model inference and training, managing data across GPU memory, local memory, distributed memory/disk, and remote storage (HDFS) to create an integrated tiered system.
Role type
Senior IC storage systems engineer (distributed storage for ML)
Builds
Integrated storage systems for large model inference and training, unified user behavior data storage, and high-throughput low-latency distributed storage solutions.
Domain
Internet / Distributed Systems / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
C/C++/Java, distributed system architecture, storage system optimization, high-throughput low-latency system design, GPU memory management, HDFS, Rocksdb, Redis, TCP/IP tuning, IO optimization
Preferred skills
Open source contributions, Paxos/Raft consensus algorithms, distributed transaction models, OS kernel knowledge
Responsibilities
Design and implement multi-level storage systems for recommendation/advertising large model inference and training; Optimize recommendation large model KV Cache hit rates across inference frameworks and traffic scheduling; Build unified storage, IO, and near-end cache solutions for exabyte-scale user behavior data supporting training and inference.
Seniority
Senior, hands-on IC