大模型推理后台链路工程师(语音方向)(北京/深圳/上海)
Core
Design and implement the end-to-end inference backend for large-scale voice models, covering access, scheduling, inference, streaming, storage, and monitoring.
Role type
Senior IC backend infrastructure engineer (LLM inference)
Builds
Real-time voice dialogue services with optimized latency and throughput
Domain
AI / Large Language Models / Voice Technology
Deliverable
production ML models
Required skills
C++ / Go / Python, Linux system programming, network programming, concurrent programming, CUDA programming, GPU operator optimization, quantization (INT8/FP8), KV Cache management, speculative decoding, continuous batching, distributed systems, Kubernetes, microservices governance, RPC frameworks
Preferred skills
Experience with TensorRT, vLLM, Triton Inference Server, ONNX Runtime, multi-machine multi-card inference, elastic scaling, fault recovery, high availability architecture, capacity planning
Technologies
TensorRT, vLLM, Triton Inference Server, ONNX Runtime, Kubernetes, CUDA
Responsibilities
Design and implement the overall architecture for voice LLM inference backends; Build streaming inference services for real-time voice-to-voice interaction; Optimize GPU inference engines including operator optimization, memory management, and throughput; Construct distributed inference capabilities with tensor/pipe parallelism and elastic scaling; Establish high-availability architectures with monitoring, alerting, and load testing; Implement cutting-edge technologies for voice LLMs and inference frameworks.