AI Benchmark Engineer
Core
Design and engineer benchmark tasks to evaluate coding agents on multilingual software challenges, specifically building realistic datasets and environments in Chinese.
Role type
Senior IC AI Benchmark Engineer (Multilingual)
Builds
Benchmark datasets, task environments, and deterministic verifier scripts for coding agents
Domain
AI Safety / Multilingual Software Engineering
Deliverable
production ML models
Required skills
Python, shell scripting, data processing, CLI-based development, multilingual text-processing (encoding, Unicode, locale), coding agent familiarity
Preferred skills
Native/near-native Chinese fluency, deep knowledge of Chinese grammar/register, experience with bidirectional/RTL handling, font fallbacks, typography
Technologies
Python, Shell, CLI tools
Responsibilities
Design Terminal-Bench tasks for coding agents, build Chinese datasets and task environments, identify AI failure points via native prompting, develop reference implementations and verifier scripts, analyze execution logs and calibrate task difficulty, participate in human review and automated quality-control processes
Seniority
Senior, hands-on IC