Senior Research Engineer (Agentic Behavior)
Core
Building evaluation infrastructure, error analysis tools, and post-training pipelines to measure and improve AI coding agents' behavior on Kotlin code.
Role type
Senior Research Engineer (Agentic Behavior)
Builds
Evaluation pipelines, observability tools, simulation environments, and open-source benchmarks for AI coding agents.
Domain
AI/ML, Software Development Tools, Kotlin Ecosystem
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python engineering, data analysis at scale, end-to-end project ownership, product-aware mindset, Kotlin familiarity
Preferred skills
Post-training LLMs (SFT, RLHF, DPO, GRPO), deep learning frameworks (PyTorch, TRL, verl, Megatron), AI agent development, evaluation frameworks (Inspect AI, Promptfoo, LM-evaluation-harness), experiment tracking (Weights & Biases, MLflow, Langfuse), Kotlin ecosystem (Android, Gradle, KMP, Spring, Ktor)
Technologies
Python, SQL, Athena, PyTorch, TRL, verl, Megatron, Inspect AI, Promptfoo, LM-evaluation-harness, Weights & Biases, MLflow, Langfuse, Kotlin, Gradle, Spring, Ktor
Responsibilities
Design and implement tooling to capture, classify, and analyze errors made by AI coding agents; Build observability pipelines over agentic traces; Design, implement, and maintain evaluation pipelines for Kotlin code generation quality; Build simulation environments for measuring agents on realistic Kotlin tasks; Experiment with post-training techniques to improve model behavior on Kotlin; Design and build open-source benchmarks for AI coding agent performance.
Seniority
Senior, hands-on IC