大模型推理工程实习生(J105856)
Core
Design and develop inference deployment platforms and inference engine infrastructure for large language model (LLM) reasoning services, building high-throughput, low-latency, and highly available systems to support large-scale token inference requests.
Role type
LLM inference infrastructure engineer (intern)
Builds
High-throughput, low-latency LLM inference systems and deployment platforms
Domain
Cloud-native AI platforms and Large Language Model inference
Deliverable
production ML models
Required skills
Go, Python, Kubernetes, algorithm and data structures, LLM inference engine frameworks (vLLM, SGLang, TensorRT-LLM), cloud-native AI platform development
Preferred skills
Separated architecture, Nvidia Dynamo, LLM-D, Mooncake, LMCache, KVCache infrastructure, multi-modal inference, Prefix Caching, Speculative Decoding
Responsibilities
Develop platform components and tools; handle custom development requests; conduct LLM inference technology research
Seniority
Intern