CareerPlanSign in

高性能计算研发工程师 - Data语音

上海💼 Full-time🗓 2026-09-28

Core

Building next-generation large model inference engines to optimize multi-modal audio generation and understanding performance on GPU clusters for low-latency, high-throughput industrial deployment.

Role type

Senior IC high-performance computing engineer (GPU inference)

Builds

High-performance inference systems and optimized deployment solutions for multi-modal audio large models

Domain

AI/ML, Audio, High-Performance Computing

Deliverable

production ML models

Required skills

Python, C++, CUDA programming, vLLM framework development, model quantization, distributed inference optimization, PCIe communication optimization

Preferred skills

Triton operator development, SGLang framework experience, TileLang development, Transformer architecture expertise, sparse model optimization

Technologies

CUDA, Triton, vLLM, SGLang, C++, Python, GPU clusters

Responsibilities

Develop and optimize multi-modal audio large model inference engines on GPU clusters; Design distributed inference strategies and communication architectures; Build high-performance inference system stacks using C++/Python; Collaborate with upstream/downstream teams to resolve performance bottlenecks and support AI toolchain development.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.