大模型推理优化工程师 - Data语音
Core
Build and optimize high-performance inference engines for multimodal speech understanding and generation large models on GPU clusters to achieve low latency and high throughput.
Role type
Senior IC large model inference optimization engineer (audio/speech)
Builds
Production-grade inference systems for speech AIGC products (e.g., DouBao, TikTok, CapCut) and enterprise services via Volcengine.
Domain
AI/ML, Large Language Models, Audio/Speech Technology, GPU Computing
Deliverable
production ML models
Required skills
Python, C++, CUDA programming, vLLM framework development, model quantization, distributed inference strategies, PCIe communication architecture, Transformer architecture knowledge
Preferred skills
Triton/TileLang development experience, sparse model optimization, research publications in model acceleration
Technologies
CUDA, Triton, vLLM, SGLang, C++, Python, GPU clusters, PCIe
Responsibilities
Develop CUDA/Triton operators and upgrade vLLM/SGLang frameworks; Design distributed inference strategies and high-concurrency architectures; Collaborate with upstream/downstream teams to analyze bottlenecks and optimize training/inference efficiency; Support AI toolchain construction and business deployment.
Seniority
Senior, hands-on IC