CareerPlanSign in

大模型Infra技术研究员-(北京)or

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Design, develop, and iterate large model inference engine architecture and multi-modal training infrastructure, optimizing performance and compute costs for high-concurrency, low-latency services.

Role type

Senior IC large model infrastructure researcher (systems)

Builds

Production-grade PD-separated inference scheduling systems and multi-modal model training infrastructure

Domain

AI Infrastructure / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

Distributed training systems, High-performance computing, MLSys, Python, C++, CUDA programming, Tensor/pipe/sequence parallelism, GPU architecture, Heterogeneous AI chips, vLLM, SGLang

Preferred skills

MLSys/SC/EuroSys/OSDI/ATC publications, CVPR/NeurIPS/ICML publications, DeepSpeed, Megatron-LM, PyTorch FSDP

Responsibilities

Design and optimize inference engine architecture for mainstream GPUs and heterogeneous AI chips; Build and optimize multi-modal model training infrastructure including memory management and cross-node communication; Implement key technologies like dynamic memory allocation, KV Cache optimization, and variable-length sequence batching; Perform full-stack performance analysis and operator optimization for training and inference; Collaborate on open-source projects like vLLM and SGLang.

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.