Director, Research - AI Evals
Core
Define quality dimensions for Figma's AI-powered experiences, build evaluation frameworks combining human and automated methods, and produce decision-ready signals for product teams.
Role type
Director, AI Evaluation Research
Builds
Evaluation frameworks, rubrics, golden datasets, and quality bars for AI features
Domain
Design collaboration software / AI/LLM evaluation
Deliverable
production ML models | dashboards & analysis
Required skills
AI/LLM product evaluation, human evaluation program design, rubric and benchmark construction, inter-rater reliability analysis, automated/model-based evaluation strategies, qualitative and quantitative research methods, stakeholder management, team leadership
Preferred skills
Automated evaluation pipeline development, regression testing, new function building, product design or data science background, Figma product familiarity
Technologies
LLM-as-judge, Braintrust, LangSmith, DeepEval
Responsibilities
Own AI evaluation methods and operations, build and maintain evaluation frameworks, partner with engineering on reproducible evaluation pipelines, produce clear readouts and dashboards, socialize shared quality definitions, manage a small team to execute AI evals
Seniority
Director, management level