大模型算法工程师(压缩与轻量化方向) - PICO
Core
Research and implement algorithms to compress and optimize large language models (LLMs) and multimodal models for efficient inference on XR devices.
Role type
Senior IC machine-learning engineer (model compression & lightweighting)
Builds
Quantized, pruned, and distilled LLMs for PICO's XR platform
Domain
Artificial Intelligence / Large Language Models / XR Hardware
Deliverable
production ML models
Required skills
Transformer architecture internals, model quantization (PTQ/QAT), pruning, knowledge distillation, Python, PyTorch, experimental analysis, mathematical foundations
Preferred skills
Low-bit quantization algorithms (AWQ, GPTQ, SmoothQuant), speculative sampling, LoGits/feature distillation, distributed training frameworks (Megatron-LM, DeepSpeed), GPU/NPU hardware optimization
Technologies
PyTorch, Python, AWQ, GPTQ, SmoothQuant, OmniQuant, vLLM, bitsandbytes, Megatron-LM, DeepSpeed, CUDA
Responsibilities
Design and implement post-training quantization and quantization-aware training schemes; Develop structured and unstructured pruning strategies for Transformer architectures; Build knowledge distillation pipelines from large teacher models to lightweight student models; Analyze accuracy degradation and propose algorithmic compensation strategies.
Seniority
Senior, hands-on IC