大模型推理优化工程师-Commercial AI
Core
Design and develop large-scale machine learning systems for high-concurrency, high-reliability, and scalable inference of large models (LLM/MLLM/Diffusion).
Role type
Senior IC large model inference optimization engineer
Builds
Distributed inference systems and optimized model deployment pipelines for commercial AI products
Domain
Commercial AI, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
C/C++, Python, Linux, distributed system design, PyTorch/TensorFlow/PaddlePaddle/Mindspore, vLLM/SGLang/TRT-LLM, model compression, hardware adaptation
Preferred skills
CUDA, vectorization, parallelization, AI compiler experience, NLP/CV/speech algorithms, domestic heterogeneous hardware tuning
Responsibilities
Design system architecture for massive ML workloads, optimize distributed inference performance and traffic scheduling, collaborate with algorithm teams on joint optimization, adapt and optimize for domestic hardware
Seniority
Senior, hands-on IC