CareerPlanSign in

内容生态研发部_ 推理性能优化工程师(J85683)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Optimize inference performance for multimodal LLMs and Diffusion Models, manage GPU cluster utilization, and develop model serving infrastructure to maintain SOTA.

Role type

Senior IC machine-learning engineer (inference optimization)

Builds

High-performance, high-concurrency, high-availability model inference services

Domain

AI/ML, Large Language Models, Video Generation

Deliverable

production ML models

Required skills

C/C++, Python, PyTorch, Transformer architecture, vLLM, SGLang, TensorRT-LLM, LightLLM, CUDA, GPU high-performance computing, parallel computing optimization, memory access optimization, low-bit computation, Docker, Linux

Preferred skills

Machine learning platform development, open-source distributed inference framework contributions

Responsibilities

Develop model serving capabilities and GPU resource scheduling functions, optimize inference performance for multimodal LLMs and Diffusion Models, collaborate on business implementation of latest research trends, solve technical challenges in high-performance and high-concurrency scenarios

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.