Product Manager - AI Evaluations Tooling
Core
Own the product strategy and roadmap for Relativity's AI evaluation platform, enabling internal teams to measure, monitor, and improve AI performance without routing questions through applied scientists.
Role type
Senior Product Manager (AI Platform/Developer Tools)
Builds
A standardized platform for offline/online AI evaluation, agentic testing systems, and team-facing quality dashboards.
Domain
Legal Technology (eDiscovery) + Artificial Intelligence (LLMs, Agentic Systems)
Required skills
Product strategy and roadmap ownership, User discovery with technical audiences, Build-vs-buy decision making, LLM evaluation methodologies (rubrics, LLM-as-judge), Agentic system testing, Stakeholder management with legal experts, Agile practices
Preferred skills
Experience with LLM-based products, Direct experience with agentic systems, Platform or developer-tools PM background
Technologies
LLMs, Agentic systems, Tracing infrastructure, Commercial eval tooling
Responsibilities
Define product vision for evaluation pillars (offline pipeline, online monitoring, agentic testing), Orchestrate agentic testing system development, Design SME authoring workflows for legal experts, Own quality dashboards for AI product teams, Lead build-vs-buy evaluations for evaluation tooling, Partner on judge calibration for LLM-as-judge
Seniority
Senior, hands-on IC