大模型代码评测专家 - AI数据与安全
Core
Design and execute evaluation benchmarks for large language models (LLMs) focusing on code generation, repair, refactoring, and agentic workflows to support model iteration.
Role type
Senior IC LLM evaluation engineer (code)
Builds
Automated evaluation tools, internal benchmark datasets, and quantitative analysis reports for model performance.
Domain
Artificial Intelligence / Large Language Models / Software Engineering
Deliverable
production ML models
Required skills
Python, LLM evaluation methodologies, Prompt Engineering, Rubric scoring, LLM-as-Judge, Agentic evaluation, data analysis, algorithm design
Preferred skills
Experience building platforms/tools from scratch, publishing papers in international conferences, project management, deep user experience with coding assistants
Technologies
Python, LLM frameworks, evaluation tooling
Responsibilities
Research and design internal code evaluation benchmarks; develop algorithms for automated evaluation and defect localization; define evaluation standards and metrics; generate data-driven reports for model iteration.
