大模型评估PM实习生(J105729)
Core
Build and optimize evaluation benchmarks and metrics for large language models (LLM/VLM) across general, vertical, and multimodal domains to support model iteration and customer scenario assessment.
Role type
Intern Product Manager (Large Model Evaluation)
Builds
Quantifiable evaluation benchmarks and assessment frameworks for LLMs
Domain
Artificial Intelligence / Large Language Models / NLP
Deliverable
production ML models
Required skills
LLM/VLM evaluation methodologies, benchmark construction, data analysis, Python/scripting, cross-functional collaboration
Preferred skills
Experience with open-source benchmarks, independent research capabilities, engineering implementation of evaluation strategies
Technologies
Python, LLM frameworks, benchmark datasets
Responsibilities
Define quantifiable evaluation standards for customer scenarios with GTM teams; Build and optimize proprietary benchmarks for stability and fairness; Adapt open-source benchmarks to full evaluation workflows; Collaborate with training, product, and algorithm teams to align evaluation systems with model iterations.