上海-AI异构计算工程师(J101238)
Core
Optimize and deploy large language models (text, image, voice) for business applications, focusing on distributed training frameworks and inference engine acceleration.
Role type
Senior IC AI Heterogeneous Computing Engineer
Builds
Distributed training frameworks for LLMs; optimized inference engines for high-traffic products
Domain
Artificial Intelligence / Large Language Models / Heterogeneous Computing
Deliverable
production ML models
Required skills
C/C++/Python, CUDA programming, operator optimization, distributed training frameworks (Megatron, DeepSpeed), inference frameworks (vLLM, SGlang), model compression techniques (SFT, RL, knowledge distillation), GPU cluster performance tuning
Preferred skills
Experience with trillion-parameter models, Agentic RL toolchains, speculative decoding, KV cache optimization
Technologies
PyTorch, Megatron, DeepSpeed, vLLM, SGlang, CUDA
Responsibilities
Build and optimize distributed training frameworks for various LLMs; optimize inference engines to reduce TTFT and increase throughput; develop reinforcement learning toolchains for efficient iteration; reduce deployment costs for high-traffic products
Seniority
Senior, hands-on IC