大模型Infra技术研究员-(北京)or
Core
Design, develop, and iterate large model inference engine architecture and multi-modal training infrastructure, optimizing performance and compute costs for high-concurrency, low-latency services.
Role type
Senior IC large model infrastructure researcher (systems)
Builds
Production-grade PD-separated inference scheduling systems and multi-modal model training infrastructure
Domain
AI Infrastructure / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed training systems, High-performance computing, MLSys, Python, C++, CUDA programming, Tensor/pipe/sequence parallelism, GPU architecture, Heterogeneous AI chips, vLLM, SGLang
Preferred skills
MLSys/SC/EuroSys/OSDI/ATC publications, CVPR/NeurIPS/ICML publications, DeepSpeed, Megatron-LM, PyTorch FSDP
Responsibilities
Design and optimize inference engine architecture for mainstream GPUs and heterogeneous AI chips; Build and optimize multi-modal model training infrastructure including memory management and cross-node communication; Implement key technologies like dynamic memory allocation, KV Cache optimization, and variable-length sequence batching; Perform full-stack performance analysis and operator optimization for training and inference; Collaborate on open-source projects like vLLM and SGLang.