Human Data Evals Lead (Remote/US/LATAM)
Core
Owns data initiatives and proposals to AI labs, designing sample packages and benchmarks to demonstrate model capability and convert pilots into production engagements.
Role type
Lead, Human Data Evaluation & Benchmarking
Builds
Frontier-grade data proposals, sample packages, and benchmarks for reasoning, coding, agents, and tool use.
Domain
Artificial Intelligence / Large Language Model Evaluation
Deliverable
production ML models
Required skills
AI/ML data proposal development, LLM benchmarking, code-model evaluation, expert recruitment and calibration, pilot delivery management, quality control (QC) process design, stakeholder management with AI labs
Preferred skills
Spanish fluency
Technologies
LLM evaluation frameworks, benchmarking tools, multi-model evaluation systems
Responsibilities
Develop data proposals and sample packages that meet frontier-grade standards; recruit, brief, and calibrate subject-matter experts; manage end-to-end pilot delivery including scoping, staffing, and QC; serve as primary contact for AI lab partners; establish and maintain rigorous quality standards for evaluation tasks.
Seniority
Lead, hands-on IC with strategic ownership