Agent数据&评测工程师/专家-Dev Infra
Core
Design and build automated evaluation infrastructure, develop evaluation agents, and construct high-quality datasets and benchmarks to measure LLM and Agent capabilities across coding and personal assistant scenarios.
Role type
Senior IC LLM evaluation and data infrastructure engineer
Builds
Automated evaluation infrastructure, evaluation agents, and large-scale high-quality datasets/benchmarks
Domain
Large Language Models (LLM) and Agent development
Deliverable
production ML models
Required skills
C/C++/Go/Python, data structures and algorithms, LLM evaluation, data engineering, benchmark construction
Preferred skills
Agent development, academic publications, innovative evaluation methods
Technologies
LLM frameworks, data parsing tools, benchmarking platforms
Responsibilities
Develop evaluation agents and automate testing workflows; construct and curate large-scale datasets for benchmarking; analyze evaluation data to drive algorithmic improvements; explore new evaluation methodologies and industry trends.
Seniority
Senior, hands-on IC
