CareerPlanSign in

AI智能体评测高级工程师

Shenzhen, China💼 Full-time🗓 2026-09-28

Core

Design and evolve evaluation systems for Tencent's self-developed AI products (WorkBuddy, CodeBuddy) and LLM/Agent systems, ensuring quality and performance improvements.

Role type

Senior AI Evaluation Engineer (LLM & Agent Systems)

Builds

Evaluation benchmarks, automated testing frameworks, and quality assurance pipelines for AI products.

Domain

Artificial Intelligence / Large Language Models / Agent Systems

Deliverable

production ML models

Required skills

Python, LLM evaluation methodologies, Transformer architecture, PyTorch, Agent development concepts (ReAct, Function Calling, Tool Use, Planning), data analysis

Preferred skills

Deep learning frameworks (TensorFlow, JAX), automated testing framework development

Responsibilities

Design evaluation systems for LLMs and Agent tasks, track industry benchmarks (SWE-bench, HumanEval, MMLU), build evaluation datasets and executors, establish evaluation standards and processes.

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.