AI Benchmark Engineer
Core
Design and engineer high-signal benchmark tasks for evaluating coding agents on multilingual software challenges.
Role type
Senior IC AI Benchmark Engineer (multilingual evaluation)
Builds
Realistic task environments, datasets, files, native-language prompts, reference implementations, and deterministic verifier scripts for AI coding agents.
Domain
Artificial Intelligence / Software Engineering / Multilingual Text Processing
Deliverable
production ML models
Required skills
Python, shell scripting, data processing, Terminal/CLI development workflows, coding agents, multilingual text processing (encoding, decoding, Unicode normalization, locale conventions, text I/O, toolchain interoperability), bidirectional/RTL handling, font fallbacks, rendering typography
Preferred skills
Experience with Haiku, Sonnet, and Opus models, human review and calibration of benchmarks, automated LLM-based quality checks, audit processes
Technologies
Python, shell scripting, Terminal-Bench configurations
Responsibilities
Design high-signal Terminal-Bench tasks, build realistic task environments and datasets, create native-language prompts, identify failure points in AI systems, develop reference implementations and verifier scripts, analyze execution logs, calibrate task difficulty, participate in human and automated quality checks, ensure benchmark fairness and integrity
Seniority
Senior, hands-on IC