CareerPlanSign in

Agent策略产品(评测方向) - 豆包

北京💼 Full-time🗓 2026-09-28

Core

Design and maintain evaluation systems for AI Agents, focusing on building high-quality benchmarks, automated evaluation pipelines, and analyzing results to drive model and product optimization.

Role type

Senior Product Manager (Agent Evaluation & Benchmarking)

Builds

Scalable automated evaluation frameworks, industry-specific benchmarks, and data production pipelines for Agent testing.

Domain

Artificial Intelligence / Large Language Models / Agent Systems

Deliverable

production ML models

Required skills

Agent evaluation methodology, benchmark design, data pipeline construction, Python, SQL, LLM-as-a-Judge, rubric creation, automated scoring, stakeholder management

Preferred skills

Complex industry domain expertise (finance, healthcare, legal, etc.), end-to-end Agent development experience, SFT/RL/Reward design knowledge, enterprise client collaboration

Technologies

Python, SQL, LLM-as-a-Judge, automated verification tools

Responsibilities

Define evaluation metrics and acceptance criteria for Agent scenarios; design and maintain high-quality evaluation datasets and benchmarks; build automated evaluation pipelines integrating rule-based and LLM-based judges; analyze evaluation failures to identify capability gaps; collaborate with algorithm and product teams to implement optimization suggestions.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.