Senior AI Evals Engineer (AI Quality & Safety) - Libra-Legal AI Assistant (m/w/d)
Core
Building and owning the evaluation platform (datasets, rubrics, judges, harnesses) to measure quality and safety for a legal AI assistant, enabling automated optimization and ensuring guardrails against hallucinations and leakage.
Role type
Senior IC AI Evals Engineer (Quality & Safety)
Builds
Self-service eval platform, automated optimization pipelines, and safety guardrails for a legal AI product.
Domain
Legal technology + Generative AI
Deliverable
production ML models | product features
Required skills
Python (FastAPI, pandas, numpy), LLM evaluation design (datasets, rubrics, LLM-as-judge), automated optimization (GEPA, DSPy), security and data-privacy practice, system architecture
Preferred skills
Automated prompt/pipeline optimization, embedding clustering, data science toolkit
Technologies
Python, FastAPI, Langfuse, GEPA, DSPy
Responsibilities
Build and own the eval platform including datasets, judges, and harnesses; Map quality landscape across different legal tasks and jurisdictions; Design rubrics and standards for LLM-as-judge; Deep-dive into results to identify failure levers and automate fixes; Unhobble the optimizer by instrumenting prompts and hyperparameters for automated search; Optimize for cost-efficiency (quality per euro); Design and red-team safety guardrails against hallucinations and prompt injection.
Seniority
Senior, hands-on IC