大模型推理存储系统专家 - Seed Model
Core
Design and develop machine learning system storage components for large model inference scenarios, optimizing data I/O performance and managing multi-level storage to improve core metrics like TTFT and TBT.
Role type
Senior IC machine learning systems engineer (storage)
Builds
Multi-level storage systems integrating GPU memory, local memory, distributed memory, and remote storage (HDFS/Object Storage) for large model inference
Domain
AI/ML Infrastructure, Distributed Systems, Cloud Native
Deliverable
production ML models
Required skills
C++, Go, Python, Linux, Kubernetes, Distributed Systems, NVLink, RDMA, GPU Direct, KV Cache optimization
Preferred skills
vLLM, SGLang, PyTorch, Alluxio, JuiceFS, GooseFS, JindoFS, OSDI/SOSP/FAST publications
Responsibilities
Design and implement multi-level storage systems for large model inference; Optimize KV Cache hit rates and data read performance; Develop efficient data access interfaces for inference frameworks; Manage storage systems in Kubernetes environments; Build multi-datacenter and multi-cloud disaster recovery systems.
Seniority
Senior, hands-on IC