Research Engineer (Agentic Behavior – Kotlin AI Value Stream) Amsterdam, Netherlands
Core
Building evaluation infrastructure, error analysis tools, and post-training pipelines to measure and improve AI coding agents' behavior on Kotlin code.
Role type
Research Engineer (Agentic Behavior)
Builds
Evaluation pipelines, observability tools, simulation environments, and open-source benchmarks for AI coding agents.
Domain
AI/ML, Large Language Models, Kotlin Ecosystem
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python engineering, data analysis at scale, end-to-end project ownership, product-aware mindset, Kotlin familiarity
Preferred skills
Post-training LLMs (SFT, RLHF, DPO, GRPO), deep learning frameworks (PyTorch), AI agent development, experiment tracking, open-source contributions
Technologies
PyTorch, TRL, verl, Megatron, Inspect AI, Promptfoo, LM-evaluation-harness, Weights & Biases, MLflow, Langfuse, Spring, Ktor, Gradle, KMP, Android
Responsibilities
Design and implement tooling to capture and analyze AI agent errors; Build evaluation pipelines measuring code generation quality; Experiment with post-training techniques to improve model behavior; Build public open-source benchmarks for Kotlin tasks.
Seniority
Mid-Senior, hands-on IC