CareerPlanSign in

大模型代码评测专家 - AI数据与安全

北京💼 Full-time🗓 2026-09-28

Core

Design and execute evaluation benchmarks for large language models (LLMs) focusing on code generation, repair, refactoring, and agentic workflows to support model iteration.

Role type

Senior IC LLM evaluation engineer (code)

Builds

Automated evaluation tools, internal benchmark datasets, and quantitative analysis reports for model performance.

Domain

Artificial Intelligence / Large Language Models / Software Engineering

Deliverable

production ML models

Required skills

Python, LLM evaluation methodologies, Prompt Engineering, Rubric scoring, LLM-as-Judge, Agentic evaluation, data analysis, algorithm design

Preferred skills

Experience building platforms/tools from scratch, publishing papers in international conferences, project management, deep user experience with coding assistants

Technologies

Python, LLM frameworks, evaluation tooling

Responsibilities

Research and design internal code evaluation benchmarks; develop algorithms for automated evaluation and defect localization; define evaluation standards and metrics; generate data-driven reports for model iteration.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.