Member of Technical Staff (Data Scientist, Evals)
Core
Build specialized automated evaluation pipelines and methods to assess answer quality for an LLM-first search engine, covering search-based answers and tool calls.
Role type
Senior IC data scientist (LLM evaluation)
Builds
Automated evaluation pipelines, evaluation sets, and VLM-based rendering solutions for Perplexity's search products
Domain
Search engines, LLMs, and agentic workflows
Deliverable
production ML models
Required skills
Python, SQL, AWS, Databricks, agentic coding workflows, LLM-as-a-judge setups, defining evaluation metrics, building ground truth datasets
Preferred skills
Experience with LLMs at scale, customer-facing web products, research background, applying research methods to real-world ML problems
Technologies
Python, SQL, AWS, Databricks, VLM
Responsibilities
Architect and maintain automated evaluation pipelines to assess answer quality; Design evaluation sets and methods to measure the impact of tool calls on answer quality; Develop VLM-based solutions to programmatically evaluate answer rendering; Continuously review and adapt public benchmarks and academic evaluations; Operate within a small team to shape product changes based on evaluation metrics
Seniority
Senior, hands-on IC