CareerPlanSign in

Stage de fin d’études — LLM-as-a-Judge : évaluation automatique, fiabilité et qualité des modèles d’IA (H/F)

Puteaux, IDF, fr💼 Full-time🗓 2026-09-17 → 2026-09-25

Core

Designing an automatic evaluation framework based on LLM-as-a-Judge to assess the reliability, bias, and quality of generative AI and agent solutions.

Role type

Intern (Stage de fin d'études) specializing in AI evaluation and LLM-as-a-Judge protocols

Builds

Evaluation frameworks, test suites, scoring grids, and monitoring integration for GenAI and agent projects

Domain

Artificial Intelligence, Generative AI, LLM Evaluation, Data & AI Consulting

Deliverable

production ML models | product features

Required skills

Python, LLM frameworks (API calls, prompt engineering, RAG), methodological rigor, experimental mindset

Preferred skills

Data science, AI strategy, digital transformation

Technologies

Python, LLM frameworks, RAG, Microsoft Fabric, Mistral, Dataiku, DataCamp

Responsibilities

Design automatic evaluation protocols for GenAI and agent solutions, build test suites and scoring grids, measure evaluation reliability and analyze judge model biases, integrate evaluation into deployment chains and monitoring, formalize a reusable evaluation framework

Seniority

Intern (final year engineering or business school student)

Sourced via smartrecruiters · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.