CareerPlanSign in

Agent评测专家(To B方向) - AI数据与安全

北京💼 Full-time🗓 2026-09-28

Core

Design and execute evaluation frameworks for AI Agents and Coding models, transforming abstract capability requirements into observable test cases and scalable data production pipelines.

Role type

Senior AI Model Evaluation Engineer (Agent/Coding)

Builds

Scalable evaluation datasets, automated assessment pipelines, and quality metrics for AI Agents and Coding models.

Domain

Artificial Intelligence / Large Language Models / Agent Systems

Deliverable

production ML models

Required skills

LLM/Agent architecture understanding, Python programming, test case design, data pipeline engineering, project management, root cause analysis

Preferred skills

Experience with automated evaluation tools, ability to define expert personas for data production, proficiency in analyzing model failure modes

Technologies

Python, LLMs, Agent frameworks, Evaluation platforms

Responsibilities

Define executable evaluation schemes for Agent and Coding models; Organize external experts to produce standardized evaluation data; Engineer automated assessment workflows and integrate them into internal platforms; Collaborate with research and product teams to drive model iteration based on evaluation data.

Sourced via bytedance · Listed on CareerPlan, which tracks 878,000+ jobs from 20+ sources.