CareerPlanSign in

混元LLM大模型量化压缩高级算法工程师(北京/深圳/上海)

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Research and develop LLM compression and acceleration algorithms to optimize inference performance and reduce model size.

Role type

Senior IC machine-learning engineer (LLM compression)

Builds

Optimized LLM inference solutions for business deployment

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

PyTorch, Python, deep learning fundamentals, LLM compression algorithms, hardware-aware optimization, sparse attention, KV-Cache compression, model pruning, quantization (FP8/INT8/NVFP4), speculative sampling, long-context optimization

Preferred skills

Top-tier conference publications, experience with hardware co-optimization

Technologies

PyTorch, FP8, INT8, NVFP4, Sparse Attention, KV-Cache

Responsibilities

Design and implement LLM compression algorithms (quantization, pruning, sparsity); Optimize inference speed for prefill and RL scenarios; Develop tools for end-to-end model compression and deployment; Analyze performance bottlenecks and customize optimization strategies; Research and publish on cutting-edge compression techniques.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.