CareerPlanSign in

资深大语言模型评测研究员-TikTok

北京💼 Full-time🗓 2026-09-28

Core

Design and execute evaluation frameworks for large language models (LLMs) to define ideal response states and improve user experience on the Tako platform.

Role type

Senior LLM evaluation researcher (AI product quality)

Builds

Quantitative and qualitative evaluation metrics, benchmark datasets, and automated assessment pipelines for LLMs.

Domain

Artificial Intelligence / Large Language Models / User Research

Deliverable

production ML models

Required skills

LLM evaluation methodology, user research (qualitative & quantitative), data analysis, benchmark design, stakeholder collaboration, project management

Preferred skills

Experience with AI ideal-state assessment mechanisms, international operations coordination

Technologies

LLM automation tools, benchmarking frameworks

Responsibilities

Define evaluation metrics using internal expert reviews, crowdsourcing, and automated LLM assessments; Maintain and curate evaluation datasets and execute routine benchmark tests; Collaborate with international operations teams to implement evaluation plans.

Sourced via bytedance · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.