AI Applied Scientist
Core
Define, measure, and improve the accuracy of an AI shopping agent through evaluation science, LLM judge fine-tuning, and rigorous experimentation.
Role type
Senior Applied Scientist (AI Evaluation & LLM Systems)
Builds
Automated evaluation frameworks, LLM judge models, accuracy metrics, and performance dashboards for an AI shopping agent.
Domain
E-commerce, AI Agents, Large Language Models (LLMs)
Deliverable
production ML models | dashboards & analysis
Required skills
LLM evaluation, fine-tuning, prompt engineering, A/B testing, causal inference, statistical rigor, ambiguity navigation, technical communication
Preferred skills
RAG, RLHF, multimodal evaluation, conversational understanding, personalization, ranking systems
Technologies
LLMs, RAG, fine-tuning frameworks, evaluation pipelines
Responsibilities
Define and evolve accuracy metrics across retrieval, ranking, and recommendations; design and run experiments to measure improvements; build and maintain evaluation datasets and benchmarks; improve LLM judges via prompting and fine-tuning; translate product questions into measurable hypotheses; identify failure modes and drive data-driven improvements; make agent performance visible and actionable.
Seniority
Senior, hands-on IC with potential for leadership
