AI Benchmark Engineer | Native Language Specialist - German - Remote
Core
Design, build, and validate rigorous evaluation benchmarks for large language models to test multilingual robustness in terminal workflows and software challenges.
Role type
Native language specialist AI benchmark engineer
Builds
Verifiable evaluation suite of Terminal-Bench tasks for multilingual software challenges
Domain
AI safety and evaluation, multilingual text processing, terminal/CLI development
Deliverable
production ML models
Required skills
Python, shell scripting, data processing, terminal/CLI development, coding agents, prompt engineering, native language fluency, encoding/decoding robustness, Unicode normalization, locale-dependent conventions
Preferred skills
experience at leading technology companies, top-tier engineering university graduation, high English proficiency, bidirectional/RTL handling, font fallbacks
Technologies
Python, shell scripting, Terminal-Bench, Haiku, Sonnet, Opus
Responsibilities
Evaluate coding agents, create realistic task environments in native language, find failure points in AI performance, support development of robust reference implementations, write deterministic verifier scripts, analyze execution logs and calibrate task difficulty, participate in human quality control processes
Seniority
Mid-level, hands-on IC