大模型存储研发Leader/架构师
Core
Design and develop specialized storage systems for large language models (LLMs), building an integrated hierarchical architecture for training and inference scenarios.
Role type
Senior IC storage systems architect (LLM infrastructure)
Builds
High-performance storage systems for LLM training and inference frameworks
Domain
Artificial Intelligence / Distributed Systems / Storage Infrastructure
Deliverable
production ML models
Required skills
Distributed systems design, C/C++/Go/Python, Linux, Data structures and algorithms, High-speed interconnects (CXL, RDMA, GPU Direct), IO optimization, High availability design
Preferred skills
Experience with large-scale distributed systems, Performance tuning, System architecture
Technologies
CXL, RDMA, GPU Direct, Linux, C, C++, Go, Python
Responsibilities
Design hierarchical storage architectures for LLM training and inference; Optimize IO paths and performance metrics (TTFT, TBT, throughput) for inference; Ensure stability and throughput for large-scale multi-GPU training scenarios; Analyze data flow characteristics to resolve bottlenecks and consistency issues.