内容生态研发部_ 推理性能优化工程师(J85683)
Core
Optimize inference performance for multimodal LLMs and Diffusion Models, manage GPU cluster utilization, and develop model serving infrastructure to maintain SOTA.
Role type
Senior IC machine-learning engineer (inference optimization)
Builds
High-performance, high-concurrency, high-availability model inference services
Domain
AI/ML, Large Language Models, Video Generation
Deliverable
production ML models
Required skills
C/C++, Python, PyTorch, Transformer architecture, vLLM, SGLang, TensorRT-LLM, LightLLM, CUDA, GPU high-performance computing, parallel computing optimization, memory access optimization, low-bit computation, Docker, Linux
Preferred skills
Machine learning platform development, open-source distributed inference framework contributions
Responsibilities
Develop model serving capabilities and GPU resource scheduling functions, optimize inference performance for multimodal LLMs and Diffusion Models, collaborate on business implementation of latest research trends, solve technical challenges in high-performance and high-concurrency scenarios
Seniority
Senior, hands-on IC