AI Benchmark Engineer | Native Language Specialist - Japanese - Remote
Core
Design, build, and validate rigorous evaluation benchmarks for large language models to test multilingual robustness in terminal workflows and coding agents.
Role type
Native language specialist AI benchmark engineer
Builds
Verifiable evaluation suite tasks (Terminal-Bench) for multilingual software challenges
Domain
AI safety evaluation, multilingual text processing, terminal/CLI development
Deliverable
production ML models
Required skills
Python, shell scripting, data processing, terminal/CLI development, coding agents, native Japanese fluency, prompt engineering
Preferred skills
experience at leading technology companies, top-tier engineering university graduation, deep understanding of multilingual text processing pitfalls (encoding, Unicode, locale conventions)
Technologies
Python, shell scripting, Terminal-Bench configurations, model tiers (Haiku, Sonnet, Opus)
Responsibilities
Create high-signal task environments using native language datasets, find failure points in AI performance in native language, write deterministic verifier scripts, analyze execution logs and calibrate task difficulty, participate in 4-layer human quality control process
Seniority
Mid-level, hands-on IC