CareerPlanGet AI match score →

AI Benchmark Engineer | Native Language Specialist - German - Remote

🌐 Remote💼 Full-time🗓 2026-06-25

Core

Design, build, and validate rigorous evaluation benchmarks for large language models to test multilingual robustness in terminal workflows and software challenges.

Role type

Native language specialist AI benchmark engineer

Builds

Verifiable evaluation suite of Terminal-Bench tasks for multilingual software challenges

Domain

AI safety and evaluation, multilingual text processing, terminal/CLI development

Deliverable

production ML models

Required skills

Python, shell scripting, data processing, terminal/CLI development, coding agents, prompt engineering, native language fluency, encoding/decoding robustness, Unicode normalization, locale-dependent conventions

Preferred skills

experience at leading technology companies, top-tier engineering university graduation, high English proficiency, bidirectional/RTL handling, font fallbacks

Technologies

Python, shell scripting, Terminal-Bench, Haiku, Sonnet, Opus

Responsibilities

Evaluate coding agents, create realistic task environments in native language, find failure points in AI performance, support development of robust reference implementations, write deterministic verifier scripts, analyze execution logs and calibrate task difficulty, participate in human quality control processes

Seniority

Mid-level, hands-on IC

Rewrite
## About the Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity ## What You'll Deliver - Task Engineering: Evaluating Coding Agents. - Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. - Prompting & Translation: finding failure points where AI does not work, in your native language - Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary). - Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus). - Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity. ## Qualifications - Experience: 1+ years of industry experience in software or prompt engineering. - Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities. - Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency. - Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing. - Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents. ## Domain Expertise - Deep technical understanding of multilingual text processing pitfalls, including: - Encoding/decoding robustness and Unicode normalization. - Locale-dependent conventions (collation, casing, non-Gregorian dates). - Text I/O, toolchain interoperability, and safe string operations. - (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts. ## Why Collaborate with Lilt? - Your schedule, your rules. As an independent contractor, work when you want, as much or as little as you want. No fixed hours, no check-ins, no micromanaging. - Get paid quickly and fairly. We respect your time and your expertise. Competitive rates, prompt payments, no chasing
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗