CareerPlanGet AI match score →

Software Engineer II - Voice Agent

China, Beijing, Beijing💼 Full-time🗓 2026-07-17 → 2026-07-31

Core

Implement and evolve voice agent capabilities including real-time speech processing, multimodal orchestration, and talking avatar technologies.

Role type

Software Engineer II - Voice Agent

Builds

Voice agents, speech services, avatar components, SDKs, and end-to-end applications

Domain

Generative AI, Voice AI, Agentic Systems

Deliverable

production ML models | product features

Required skills

Real-time speech processing, multimodal orchestration, generative AI integration, system latency optimization, scalability engineering, robust system integration, C/C++/C#/Java/JavaScript/Python, machine learning-powered systems, speech processing, NLP, computer vision, real-time systems, generative AI technologies, model integration, large-scale distributed systems

Preferred skills

Model development, diffusion models, autoregressive approaches, optimization

Technologies

C, C++, C#, Java, JavaScript, Python

Responsibilities

Implement and evolve voice agent capabilities, Develop and integrate talking avatar technologies, Apply and adapt modern generative AI technologies, Collaborate with cross-functional teams to integrate components, Continuously improve system latency, quality, scalability, and reliability, Contribute to engineering excellence through design and code reviews

Seniority

Mid-level, hands-on IC

Rewrite
## About the role - Implement and evolve voice agent capabilities, including real‑time speech processing, multimodal orchestration, and agent runtime integration - Develop and integrate talking avatar technologies, supporting both zero‑shot and customized experiences - Apply and adapt modern generative AI technologies (including diffusion and language modeling) where appropriate, with an emphasis on engineering robustness and system integration - Collaborate with cross‑functional teams to integrate voice agents, speech services, and avatar components into platforms, SDKs, and end‑to‑end applications - Continuously improve system latency, quality, scalability, and reliability in real‑world deployments - Stay current with industry and research advancements in voice AI, agentic systems, and generative technologies, and translate them into practical solutions - Contribute to engineering excellence through design reviews, code reviews, and shared best practices ## Requirements - Bachelor's Degree in Computer Science or a related technical field AND 2+ years of professional software engineering experience - Strong coding skills in one or more languages such as C, C++, C#, Java, JavaScript, or Python - Experience working with machine learning-powered systems, platforms, or services (model development experience is a plus but not required) - Background in speech processing, natural language processing (NLP), computer vision, or real‑time systems - Familiarity with generative AI technologies (e.g., diffusion, autoregressive approaches), model integration, or optimization - Experience building or integrating voice agents, AI services, or large‑scale distributed systems - Strong problem‑solving skills with a systems‑thinking mindset - Effective communication skills and experience working in cross‑functional teams
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗