浏览器数据后处理算法工程师
Core
Develop and optimize browser-side/cloud collaborative capabilities for web text cleaning, content extraction, intelligent slicing, summarization, and metadata extraction; explore small model (1B~7B) technical limits for efficient deployment on edge/low-power devices.
Role type
Senior IC NLP/LLM algorithm engineer (small model post-training & optimization)
Builds
Browser-side and cloud collaborative text processing pipelines; optimized small language models for edge deployment
Domain
Internet / Browser technology / NLP / LLM
Deliverable
production ML models
Required skills
Small model post-training (SFT, DPO, RLHF), Efficient Fine-Tuning (LoRA, QLoRA), Distributed training frameworks (DeepSpeed, Megatron-LM, Ray), Large-scale web data cleaning & parsing (HTML/DOM), Model quantization & compression, Python/C++ coding, PyTorch ecosystem
Preferred skills
Experience with Llama, Qwen, Phi, Gemma model families, Synthesizing Data, Reward Model construction, Long context extension
Technologies
PyTorch, DeepSpeed, Megatron-LM, Ray, LoRA, QLoRA, C++, Python
Responsibilities
Develop and optimize web text cleaning, content extraction, slicing, and summarization for browser/cloud scenarios; Iterate on small model architectures for efficient fine-tuning and low-power deployment; Build high-quality post-processing corpus using synthetic data and reward models
Seniority
Senior, hands-on IC