CareerPlanSign in

Research Engineer (Evals)

Paris💼 Full-time🗓 2026-06-28 → 2026-09-26

Core

Build and maintain internal benchmark suites for single/multi-turn content and agentic guardrails, and study agent behaviors in the wild.

Role type

Research Engineer (Evals)

Builds

Internal benchmark suites, evals for flagship models, and tooling for studying agent reliability.

Domain

AI Safety, LLM Evaluation, Agentic Systems

Deliverable

production ML models | product features

Required skills

LLM benchmark construction, synthetic data generation for post-training, Python production code development, efficient LLM inference orchestration, frontier model usage, coding agents

Preferred skills

Automated red-teaming, experience with agentic scaffolds, reward-model/safety benchmark knowledge, published papers in evals/safety-evaluation

Technologies

Python, LLMs, coding agents

Responsibilities

Own and maintain internal benchmark suite, build benchmarks distinguishing specific model capabilities, work with product team on flagship model evals, build benchmarks for new research features, adapt evals to new verticals, study and quantify realistic agentic/LLM failure modes

Seniority

Mid-Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.