AI评测专家 - 飞书
Core
Build and maintain the evaluation system for AI Agents, covering capabilities, business impact, tool usage, task execution, stability, safety, and user experience.
Role type
Senior AI Agent Evaluation Engineer
Builds
Automated evaluation pipelines, high-quality datasets, and monitoring systems for AI Agents
Domain
Enterprise AI / Large Language Models / AI Agents
Deliverable
production ML models
Required skills
LLM evaluation, algorithm evaluation, data analysis, test development, automated evaluation platform construction, complex task E2E evaluation, tool use evaluation, Computer Use Agent evaluation, Python, SQL, structured analysis, report writing
Preferred skills
RAG, Prompt Engineering, model inference principles, LLM-as-a-Judge, failure case root cause analysis
Technologies
Python, SQL, LLM-as-a-Judge
Sourced via bytedance · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.