Machine Learning Engineering Intern
Core
Building datasets and evaluation frameworks to assess the quality of AI agents on the Zendesk Explore platform.
Role type
Machine Learning Engineering Intern (Evaluation & Data Curation)
Builds
Golden datasets and LLM-as-Judge evaluation logic for AI agents
Domain
Customer Experience SaaS / AI Agents / LLM Evaluation
Deliverable
production ML models
Required skills
Python, TypeScript/JavaScript, Large Language Models (LLMs), Natural Language Processing (NLP), Data Annotation, Prompt Engineering
Preferred skills
AWS, Back-end data pipelines, Analytics industry knowledge
Technologies
Python, TypeScript, JavaScript, Braintrust platform
Responsibilities
Annotate and curate real-world conversation datasets for benchmarking AI agents; Develop and refine evaluation logic using LLM-as-Judge techniques; Compare automated evaluation scores with human feedback to iterate on prompts and logic.
Seniority
Intern