Agent策略产品(评测方向) - 豆包
Core
Design and maintain evaluation systems for AI Agents, focusing on building high-quality benchmarks, automated evaluation pipelines, and analyzing results to drive model and product optimization.
Role type
Senior Product Manager (Agent Evaluation & Benchmarking)
Builds
Scalable automated evaluation frameworks, industry-specific benchmarks, and data production pipelines for Agent testing.
Domain
Artificial Intelligence / Large Language Models / Agent Systems
Deliverable
production ML models
Required skills
Agent evaluation methodology, benchmark design, data pipeline construction, Python, SQL, LLM-as-a-Judge, rubric creation, automated scoring, stakeholder management
Preferred skills
Complex industry domain expertise (finance, healthcare, legal, etc.), end-to-end Agent development experience, SFT/RL/Reward design knowledge, enterprise client collaboration
Technologies
Python, SQL, LLM-as-a-Judge, automated verification tools
Responsibilities
Define evaluation metrics and acceptance criteria for Agent scenarios; design and maintain high-quality evaluation datasets and benchmarks; build automated evaluation pipelines integrating rule-based and LLM-based judges; analyze evaluation failures to identify capability gaps; collaborate with algorithm and product teams to implement optimization suggestions.
Seniority
Senior, hands-on IC
