AI Product Analyst (Developer Community), AI Verify Foundation
Core
Design and run offline A/B experiments to optimize evaluation pipelines for LLM benchmark tests, lead quality assurance of datasets and evaluators, and translate research insights into actionable prototypes for engineering teams.
Role type
Manager-level AI Product Analyst (Developer Community)
Builds
Open-source tools for evaluating LLM applications (Project Moonshot) and predictive ML models (AI Verify Toolkit)
Domain
Generative AI evaluation, benchmark testing, open-source AI tools
Deliverable
production ML models | product features
Required skills
experimental design, statistical analysis, Python, R, data storytelling, prototype development
Preferred skills
AI testing experience, knowledge of GenAI evaluation best practices
Responsibilities
Design and run offline A/B experiments for benchmark tests, generate realistic test cases for Singapore context, research emerging GenAI evaluation tools, educate and engage the developer community
