资深大语言模型评测研究员-TikTok
Core
Design and execute evaluation frameworks for large language models (LLMs) to define ideal response states and improve user experience on the Tako platform.
Role type
Senior LLM evaluation researcher (AI product quality)
Builds
Quantitative and qualitative evaluation metrics, benchmark datasets, and automated assessment pipelines for LLMs.
Domain
Artificial Intelligence / Large Language Models / User Research
Deliverable
production ML models
Required skills
LLM evaluation methodology, user research (qualitative & quantitative), data analysis, benchmark design, stakeholder collaboration, project management
Preferred skills
Experience with AI ideal-state assessment mechanisms, international operations coordination
Technologies
LLM automation tools, benchmarking frameworks
Responsibilities
Define evaluation metrics using internal expert reviews, crowdsourcing, and automated LLM assessments; Maintain and curate evaluation datasets and execute routine benchmark tests; Collaborate with international operations teams to implement evaluation plans.
