Research Scientist, APEX Benchmarks
Core
Design and lead the next generation of APEX benchmarks and expert-built datasets to measure frontier AI models' performance on economically valuable professional work.
Role type
Senior Research Scientist (LLM Evaluation & Benchmarking)
Builds
APEX benchmark family (investment banking, corporate law, consulting, medicine, accounting, software engineering) and associated open datasets/papers.
Domain
AI/ML evaluation, LLM benchmarking, professional services (finance, law, consulting, medicine, accounting, SWE)
Deliverable
production ML models | product features | research
Required skills
LLM evaluation, benchmark design, statistical rigor, experimental design, coding (eval harnesses), data analysis, technical writing, cross-functional collaboration
Preferred skills
Ph.D. in ML/NLP, top-tier publications (NeurIPS/ICML/ACL/ICLR), experience authoring public benchmarks, frontier lab experience, domain expertise in finance/law/consulting/accounting/medicine/SWE
Technologies
LLMs, evaluation frameworks, statistical analysis tools
Responsibilities
Design task taxonomy, difficulty calibration, and contamination controls for benchmarks; create expert-built datasets and grading rubrics; define measurement standards (confidence intervals, inter-rater agreement); partner with academic/industry collaborators; publish research (papers, datasets, talks); translate findings into business narratives; collaborate with data ops, product, and strategy teams.
Seniority
Senior, hands-on IC with strategic impact
