CareerPlanGet AI match score →

Senior AI Software Engineer - Model Evaluation (f/m/d)

Heidelberg💼 Full-time🗓 2026-04-17 → 2026-07-31

Core

Designing, implementing, and analyzing benchmarks to evaluate pre-training foundation models and determine training decisions.

Role type

Senior AI Software Engineer (Model Evaluation)

Builds

Evaluation pipelines, benchmark suites, and reporting dashboards for pre-training runs.

Domain

Generative AI, Large Language Models (LLM), Pre-training

Deliverable

production ML models

Required skills

LLM evaluation, benchmark design, dataset curation, experimental design, statistical methods, Python, PyTorch, distributed systems

Preferred skills

foundation model training, large-scale data processing, German language proficiency

Technologies

PyTorch, distributed systems

Responsibilities

Own benchmarks end-to-end from curation to analysis; Build and optimize evaluation pipelines; Design aggregation and reporting tools; Close capability gaps with product teams; Ensure rigorous assessment of German language capabilities; Correlate pre-training metrics with downstream performance

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗