Applied AI Researcher, Benchmarking
Core
Design evaluation frameworks and construct benchmarks to measure reasoning depth, interaction quality, reliability, and operational impact of AI systems.
Role type
Applied AI Researcher (Benchmarking)
Builds
Evaluation frameworks, test suites, and experimental benchmarks for intelligent systems
Domain
Applied AI, Enterprise AI Operations
Deliverable
production ML models
Required skills
Designing and running evaluations, Statistical and analytical rigor, Building with models (compound AI systems, agentic collaboration, ensembling, ReAct, graph-of-thoughts), Proven track record of research results, Strong programming and data analysis skills
Preferred skills
Experience using AI tools daily (ChatGPT, Cursor, Perplexity)
Technologies
Compound AI systems, Agentic collaboration, Ensembling, ReAct, Graph-of-thoughts
Responsibilities
Design evaluation frameworks capturing reasoning depth and operational impact; Construct benchmarks reflecting real-world complexity; Explore new paradigms for evaluating intelligent systems (adversarial robustness, longitudinal tracking, human-in-the-loop); Investigate how metrics shape model behavior; Establish rigorous methodologies for quantifying emergent capability
Seniority
Senior, hands-on IC