高性能计算研发工程师-Seed
Core
Develop high-performance inference for voice multimodal LLM scenarios and optimize inference performance for various business applications.
Role type
Senior IC high-performance computing engineer (LLM inference)
Builds
High-performance inference systems for voice multimodal LLMs
Domain
AI / Machine Learning / High-Performance Computing
Deliverable
production ML models
Required skills
C/C++, Python, CUDA, GPU optimization, parallel programming, memory optimization, low-bit computation, deep learning algorithms, LLM model structures
Preferred skills
AI engineering, quantization, sparsity, Triton, Tilelang, CuteDSL, voice/video model optimization
Responsibilities
Develop high-performance inference for voice multimodal LLM scenarios, optimize inference performance for business scenarios, follow frontier technologies in LLM model efficiency, build leading high-performance computing capabilities