CareerPlanSign in

大模型评估PM实习生(J105729)

北京市💼 Full-time🗓 2026-09-24 → 2026-09-28

Core

Build and optimize evaluation benchmarks and metrics for large language models (LLM/VLM) across general, vertical, and multimodal domains to support model iteration and customer scenario assessment.

Role type

Intern Product Manager (Large Model Evaluation)

Builds

Quantifiable evaluation benchmarks and assessment frameworks for LLMs

Domain

Artificial Intelligence / Large Language Models / NLP

Deliverable

production ML models

Required skills

LLM/VLM evaluation methodologies, benchmark construction, data analysis, Python/scripting, cross-functional collaboration

Preferred skills

Experience with open-source benchmarks, independent research capabilities, engineering implementation of evaluation strategies

Technologies

Python, LLM frameworks, benchmark datasets

Responsibilities

Define quantifiable evaluation standards for customer scenarios with GTM teams; Build and optimize proprietary benchmarks for stability and fairness; Adapt open-source benchmarks to full evaluation workflows; Collaborate with training, product, and algorithm teams to align evaluation systems with model iterations.

Sourced via baidu · Listed on CareerPlan, which tracks 833,000+ jobs from 20+ sources.