CareerPlanSign in

AI评测专家 - 飞书

北京💼 Full-time🗓 2026-09-28

Core

Build and maintain the evaluation system for AI Agents, covering capabilities, business impact, tool usage, task execution, stability, safety, and user experience.

Role type

Senior AI Agent Evaluation Engineer

Builds

Automated evaluation pipelines, high-quality datasets, and monitoring systems for AI Agents

Domain

Enterprise AI / Large Language Models / AI Agents

Deliverable

production ML models

Required skills

LLM evaluation, algorithm evaluation, data analysis, test development, automated evaluation platform construction, complex task E2E evaluation, tool use evaluation, Computer Use Agent evaluation, Python, SQL, structured analysis, report writing

Preferred skills

RAG, Prompt Engineering, model inference principles, LLM-as-a-Judge, failure case root cause analysis

Technologies

Python, SQL, LLM-as-a-Judge

Sourced via bytedance · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.