高性能计算研发工程师 - Data语音
Core
Building next-generation large model inference engines to optimize multi-modal audio generation and understanding performance on GPU clusters for low-latency, high-throughput industrial deployment.
Role type
Senior IC high-performance computing engineer (GPU inference)
Builds
High-performance inference systems and optimized deployment solutions for multi-modal audio large models
Domain
AI/ML, Audio, High-Performance Computing
Deliverable
production ML models
Required skills
Python, C++, CUDA programming, vLLM framework development, model quantization, distributed inference optimization, PCIe communication optimization
Preferred skills
Triton operator development, SGLang framework experience, TileLang development, Transformer architecture expertise, sparse model optimization
Technologies
CUDA, Triton, vLLM, SGLang, C++, Python, GPU clusters
Responsibilities
Develop and optimize multi-modal audio large model inference engines on GPU clusters; Design distributed inference strategies and communication architectures; Build high-performance inference system stacks using C++/Python; Collaborate with upstream/downstream teams to resolve performance bottlenecks and support AI toolchain development.
Seniority
Senior, hands-on IC