CareerPlanSign in

大模型推理后台链路工程师(语音方向)(北京/深圳/上海)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design and implement the end-to-end inference backend for large-scale voice models, covering access, scheduling, inference, streaming, storage, and monitoring.

Role type

Senior IC backend infrastructure engineer (LLM inference)

Builds

Real-time voice dialogue services with optimized latency and throughput

Domain

AI / Large Language Models / Voice Technology

Deliverable

production ML models

Required skills

C++ / Go / Python, Linux system programming, network programming, concurrent programming, CUDA programming, GPU operator optimization, quantization (INT8/FP8), KV Cache management, speculative decoding, continuous batching, distributed systems, Kubernetes, microservices governance, RPC frameworks

Preferred skills

Experience with TensorRT, vLLM, Triton Inference Server, ONNX Runtime, multi-machine multi-card inference, elastic scaling, fault recovery, high availability architecture, capacity planning

Technologies

TensorRT, vLLM, Triton Inference Server, ONNX Runtime, Kubernetes, CUDA

Responsibilities

Design and implement the overall architecture for voice LLM inference backends; Build streaming inference services for real-time voice-to-voice interaction; Optimize GPU inference engines including operator optimization, memory management, and throughput; Construct distributed inference capabilities with tensor/pipe parallelism and elastic scaling; Establish high-availability architectures with monitoring, alerting, and load testing; Implement cutting-edge technologies for voice LLMs and inference frameworks.

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.