混元语音大模型agentic评测(北京/上海)
Core
Design and implement end-to-end evaluation systems for complex AI agents, focusing on multi-step reasoning, tool orchestration, and long-term planning.
Role type
Senior IC AI Agent Evaluation Engineer
Builds
End-to-end evaluation pipelines and benchmarks for agentic workflows
Domain
Artificial Intelligence / Large Language Models / Agent Systems
Deliverable
production ML models
Required skills
Agent benchmark design, multi-step reasoning evaluation, tool orchestration assessment, data quality control, Python programming, statistical analysis, LLM API integration, ReAct paradigm, task planning, MCP protocols
Preferred skills
Multimodal agent evaluation, automated evaluation infrastructure, academic/industry trend analysis
Responsibilities
Design and implement evaluation task systems for complex agent scenarios; build data quality control mechanisms including plagiarism detection and difficulty calibration; collaborate with training teams to derive model optimization insights; establish automated evaluation pipelines; explore and validate new multimodal evaluation approaches.
