AI智能体评测高级工程师
Core
Design and evolve evaluation systems for Tencent's self-developed AI products (WorkBuddy, CodeBuddy) and LLM/Agent systems, ensuring quality and performance improvements.
Role type
Senior AI Evaluation Engineer (LLM & Agent Systems)
Builds
Evaluation benchmarks, automated testing frameworks, and quality assurance pipelines for AI products.
Domain
Artificial Intelligence / Large Language Models / Agent Systems
Deliverable
production ML models
Required skills
Python, LLM evaluation methodologies, Transformer architecture, PyTorch, Agent development concepts (ReAct, Function Calling, Tool Use, Planning), data analysis
Preferred skills
Deep learning frameworks (TensorFlow, JAX), automated testing framework development
Responsibilities
Design evaluation systems for LLMs and Agent tasks, track industry benchmarks (SWE-bench, HumanEval, MMLU), build evaluation datasets and executors, establish evaluation standards and processes.