Ingénieur Data Science - Génération de données synthétiques - Stage - H/F
Core
Design and improve methods for generating synthetic datasets to train small language models (SLMs) and AI agents, focusing on structured expert process knowledge.
Role type
Intern Data Scientist (Synthetic Data Generation)
Builds
Synthetic datasets, prototypes, and evaluation metrics for SLM training and AI agent development.
Domain
Generative AI, Data Engineering, Industrial Systems (Nuclear, Energy)
Deliverable
production ML models | product features
Required skills
Python development, Data APIs, Data schemas, Data validation, Data pipelines, LLM agents, Tool calling, Synthetic data, Active learning, Git, Docker
Preferred skills
RLHF, Post-training evaluation, PyTorch
Responsibilities
Analyze reference data pipelines to identify quality and coverage improvements; Design or enhance data generation methods using teacher models, counterfactual scenarios, or active sampling; Represent states, actions, tool calls, and reasoning traces; Anchor synthetic traces in expert feedback and deterministic tool results; Evaluate data quality, diversity, difficulty, and utility for downstream training; Deliver a reproducible prototype, dataset sample, and recommendations.
Seniority
Intern (End of studies)