CareerPlanSign in

AI Benchmark Engineer

🌐 Remote💼 Full-time🗓 2026-09-25 → 2026-09-26

Core

Design and engineer high-signal benchmark tasks for evaluating coding agents on multilingual software challenges.

Role type

Senior IC AI Benchmark Engineer (multilingual evaluation)

Builds

Realistic task environments, datasets, files, native-language prompts, reference implementations, and deterministic verifier scripts for AI coding agents.

Domain

Artificial Intelligence / Software Engineering / Multilingual Text Processing

Deliverable

production ML models

Required skills

Python, shell scripting, data processing, Terminal/CLI development workflows, coding agents, multilingual text processing (encoding, decoding, Unicode normalization, locale conventions, text I/O, toolchain interoperability), bidirectional/RTL handling, font fallbacks, rendering typography

Preferred skills

Experience with Haiku, Sonnet, and Opus models, human review and calibration of benchmarks, automated LLM-based quality checks, audit processes

Technologies

Python, shell scripting, Terminal-Bench configurations

Responsibilities

Design high-signal Terminal-Bench tasks, build realistic task environments and datasets, create native-language prompts, identify failure points in AI systems, develop reference implementations and verifier scripts, analyze execution logs, calibrate task difficulty, participate in human and automated quality checks, ensure benchmark fairness and integrity

Seniority

Senior, hands-on IC

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.