微信 -WeLM 大模型推理优化工程师(深圳、上海)
Core
Optimize inference performance for large language models (LLMs) by reducing latency, increasing throughput, and improving resource efficiency across various hardware platforms.
Role type
Senior IC machine-learning inference optimization engineer
Builds
High-performance, scalable LLM inference services supporting real-time and batch processing scenarios
Domain
Artificial Intelligence / Large Language Models / GPU Computing
Deliverable
production ML models
Required skills
C++, Python, CUDA programming, GPU performance tuning, PyTorch, JAX, model quantization, model sparsification, Transformer architecture, KV Cache optimization, profiling tools (Nsight)
Preferred skills
Experience with dedicated AI chips, designing inference-friendly model architectures
Technologies
PyTorch, JAX, CUDA, Nsight, GPU, AI chips
Responsibilities
Develop and apply model compression techniques; optimize inference frameworks for multiple hardware platforms; design stable and scalable inference service architectures; establish performance benchmarking frameworks; analyze performance bottlenecks and implement optimization strategies; translate cutting-edge research into production environments.
Seniority
Senior, hands-on IC