Product Manager, AI Evaluation Lead
Core
Own the end-to-end AI evaluation program to build scalable systems, data sets, and an expert network that prove and improve AI quality for construction contract intelligence.
Role type
Senior Product Manager, AI Evaluation Lead
Builds
Scalable 'Eval Ops' system, canonical evaluation data sets, ground truth rubrics, and a tiered human-in-the-loop review pipeline.
Domain
Construction technology (AECO) and AI/LLM evaluation
Deliverable
production ML models
Required skills
Program and process ownership, network-building and incentive design, AI/LLM evaluation fluency, stakeholder management, project and resource management, quality metrics definition
Preferred skills
Experience building expert networks or annotation pipelines, technical fluency with internal tooling, exposure to construction domain, B2B SaaS product experience
Technologies
LLM-as-a-judge, human-in-the-loop review, automated evaluation pipelines
Responsibilities
Define evaluation methodology and test sets, build canonical evaluation data sets with ground truth, stand up scalable eval pipelines combining automated and human review, recruit and coordinate a flexible SME network, design incentive models for contributors, partner with Product and Engineering, establish customer feedback loops, measure and report evaluation quality metrics
Seniority
Senior, hands-on IC