CareerPlanSign in

音频算法工程师-抖音直播

北京💼 Full-time🗓 2026-09-28

Core

Building low-latency audio understanding and generation models for real-time dialogue agents in Douyin live streaming, covering ASR, TTS, end-to-end speech large models, and audio classification.

Role type

Senior IC machine-learning engineer (audio)

Builds

Real-time dialogue agents for live streaming

Domain

Consumer media / Audio AI

Deliverable

production ML models

Required skills

End-to-end speech large models, ASR, TTS, audio classification, Python, C++, TensorFlow, PyTorch, Linux, data structures

Preferred skills

Publications in ICASSP, Interspeech, NIPS, ICML, ICLR, top-tier competition awards

Responsibilities

Optimize algorithms for key scenarios to build high-quality low-latency agent systems, explore boundaries of multimodal perception and interaction capabilities, and land products.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.