Research Scientist, AI Evaluation Science
Core
Formulating open problems in evaluation science, designing experiments, and publishing findings to advance the field of AI evaluation methodology.
Role type
Senior IC research scientist (AI evaluation science)
Builds
Original research methods and production-ready evaluation SDKs/APIs
Domain
Artificial Intelligence / Machine Learning / Measurement Science
Deliverable
production ML models
Required skills
Preference learning, reward modeling, calibration theory, statistical reliability, human-AI interaction methodology, experimental design, publication record in top-tier venues
Preferred skills
Measurement theory, validity frameworks, statistical learning theory, RLHF, LLM-as-judge approaches, benchmark design, agentic system evaluation
Technologies
PyTorch, JAX, TensorFlow
Responsibilities
Advance evaluation methodology through original research in preference learning, reward modeling, calibration, or validity frameworks; Publish at top-tier venues (NeurIPS, ICML, ICLR, ACL, EMNLP); Translate research into production-ready tools by partnering with platform engineers; Collaborate with measurement scientists to integrate psychometric methods; Define the team's research agenda for evaluation science
Seniority
Senior, hands-on IC