CareerPlanSign in

语音算法实习生(J84146)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Research and develop large-scale speech synthesis and understanding systems, specifically LLM-based TTS and end-to-end Speech2Speech architectures.

Role type

Research intern (Speech AI / Large Language Models)

Builds

LLM-based Text-to-Speech (TTS) and Speech2Speech systems

Domain

Artificial Intelligence / Speech Processing / Large Language Models

Deliverable

production ML models

Required skills

Python, C/C++, PyTorch, TTS algorithms, ASR algorithms, VITS, VALLE, FishSpeech, CosyVoice, LLM theory

Preferred skills

End-to-end speech modeling, multimodal alignment, model evaluation and optimization

Responsibilities

Decouple and analyze modules in latest speech synthesis and understanding solutions; Participate in R&D of speech Encodec, Decoder, and multimodal alignment modules; Evaluate and optimize speech large models

Sourced via baidu · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.