AI大模型应用能力评测专家(自动评测方向) - AI数据与安全
Core
Design and execute automated evaluation workflows for AI large language models (LLMs) to assess capabilities, risks, and performance across multi-modal and complex business scenarios.
Role type
Senior IC AI Model Evaluation Engineer (Automated Evaluation)
Builds
Automated evaluation workflows, evaluation datasets, scoring standards, and quality assurance rules for LLMs.
Domain
Artificial Intelligence / Large Language Models / Data Quality
Deliverable
production ML models
Required skills
LLM evaluation framework design, LLM as a Judge implementation, automated workflow design, data analysis, SQL, Python, structured requirement decomposition, risk identification.
Preferred skills
Experience with multi-modal data, model standardization体系建设, statistical analysis, cross-functional collaboration.
Technologies
SQL, Python, Excel, Feishu, LLM frameworks.
Responsibilities
Build vertical business scenario evaluation sets and annotation standards; Design model evaluation standard systems and SOPs; Design and iterate automated evaluation workflows; Evaluate model outputs and attribute negative cases; Analyze evaluation data and generate trend reports.
