CareerPlanSign in

大模型推理工程实习生(J105856)

北京市,上海市💼 Full-time🗓 2026-09-24 → 2026-09-28

Core

Design and develop inference deployment platforms and inference engine infrastructure for large language model (LLM) reasoning services, building high-throughput, low-latency, and highly available systems to support large-scale token inference requests.

Role type

LLM inference infrastructure engineer (intern)

Builds

High-throughput, low-latency LLM inference systems and deployment platforms

Domain

Cloud-native AI platforms and Large Language Model inference

Deliverable

production ML models

Required skills

Go, Python, Kubernetes, algorithm and data structures, LLM inference engine frameworks (vLLM, SGLang, TensorRT-LLM), cloud-native AI platform development

Preferred skills

Separated architecture, Nvidia Dynamo, LLM-D, Mooncake, LMCache, KVCache infrastructure, multi-modal inference, Prefix Caching, Speculative Decoding

Responsibilities

Develop platform components and tools; handle custom development requests; conduct LLM inference technology research

Seniority

Intern

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.