Senior Evaluation Algorithm Engineer
Core
Design and build LLM evaluation systems for dialogue and financial trading scenarios to quantify model performance and guide R&D direction.
Role type
Senior IC evaluation algorithm engineer (LLM)
Builds
Evaluation metric systems, rubrics, datasets, benchmarks, and automated evaluation platforms
Domain
AI/ML, Large Language Models, Financial Trading, Dialogue Systems
Deliverable
production ML models
Required skills
LLM evaluation methodology, rubric design, data annotation guidelines, Python, evaluation workflow automation, statistical analysis, cross-team collaboration
Preferred skills
Dialogue system evaluation, AI Agents evaluation, financial/trading LLM evaluation, RLHF, reward models, preference data
Technologies
Python
Responsibilities
Design end-to-end LLM evaluation plans and metric systems; Lead design and construction of evaluation datasets and benchmarks; Analyze model capability boundaries and failure modes; Drive automation and scaling of evaluation workflows; Collaborate with teams to translate evaluation findings into R&D directions
Seniority
Senior, hands-on IC
