AI Benchmark Engineer | Native Language Specialist - Hindi - Remote
Core
Design, build, and validate rigorous evaluation benchmarks (Terminal-Bench) to test large language models on multilingual software challenges, specifically focusing on non-English data processing and locale/encoding edge cases.
Role type
Native language specialist / AI benchmark engineer
Builds
Verifiable evaluation suites and task environments for multilingual coding agents
Domain
Artificial Intelligence / Large Language Models / Multilingual Software Engineering
Deliverable
production ML models
Required skills
Python, shell scripting, data processing, terminal/CLI development, coding agents, multilingual text processing, Unicode normalization, locale-dependent conventions
Preferred skills
Experience at leading technology companies, top-tier engineering university graduation
Technologies
Python, shell scripting, Terminal-Bench, Haiku, Sonnet, Opus
Responsibilities
Create high-signal tasks in native language, find failure points in AI models, write deterministic verifier scripts, analyze execution logs, calibrate task difficulty, participate in human quality control processes
Seniority
Senior, hands-on IC