Agent Evaluation Intern
Core
Build transferable expertise in agent evaluation engineering for a real-world LLM agent system in UA performance marketing, focusing on evaluating tool use, measuring trajectory quality, and improving reliability.
Role type
Intern, LLM agent evaluation engineer
Builds
Automated evaluation pipelines, reusable evaluation datasets, and monitoring signals for agentic AI systems
Domain
Artificial Intelligence / Large Language Models / Performance Marketing
Deliverable
production ML models
Required skills
Python, LLM agent evaluation methodology, data analysis, log and trace analysis, benchmark design, metric definition
Preferred skills
LangChain-like agents, tool calling, pytest, observability tools, research paper composition
Technologies
Python, LangChain, pytest
Responsibilities
Research state-of-the-art agentic workflow evaluation frameworks; build automated evaluation pipelines to run agent scenarios and score results; evaluate tool-use behavior and error handling; analyze agent trajectories to identify reasoning failures; design metrics for agent reliability; create reusable evaluation datasets from synthetic and real cases; support experiments comparing prompts and model providers; help build human evaluation workflows and rubrics; translate evaluation findings into better tests and guardrails
Seniority
Intern, research engineering