混元LLM大模型量化压缩高级算法工程师(北京/深圳/上海)
Core
Research and develop LLM compression and acceleration algorithms to optimize inference performance and reduce model size.
Role type
Senior IC machine-learning engineer (LLM compression)
Builds
Optimized LLM inference solutions for business deployment
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
PyTorch, Python, deep learning fundamentals, LLM compression algorithms, hardware-aware optimization, sparse attention, KV-Cache compression, model pruning, quantization (FP8/INT8/NVFP4), speculative sampling, long-context optimization
Preferred skills
Top-tier conference publications, experience with hardware co-optimization
Technologies
PyTorch, FP8, INT8, NVFP4, Sparse Attention, KV-Cache
Responsibilities
Design and implement LLM compression algorithms (quantization, pruning, sparsity); Optimize inference speed for prefill and RL scenarios; Develop tools for end-to-end model compression and deployment; Analyze performance bottlenecks and customize optimization strategies; Research and publish on cutting-edge compression techniques.
Seniority
Senior, hands-on IC
