AI Engineer (Evaluation)
Core
Designing and building evaluation frameworks and pipelines to test the capabilities of Large Language Models (LLMs), specifically focusing on multilingual, multicultural, and multimodal aspects.
Role type
AI Engineer (Evaluation)
Builds
Evaluation frameworks, pipelines, and datasets for LLMs
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
LLM evaluation experimental design, LLM inference frameworks (vLLM), Deep Learning frameworks (PyTorch), Python coding, version control (Git), data preparation and analysis, AI modelling, testing, validation, deployment
Preferred skills
Reading and understanding research papers, community engagement
Technologies
vLLM, PyTorch, Git
Responsibilities
Develop and maintain evaluation frameworks and pipelines; experiment with latest research in multilingual/multicultural/multimodal LLM evaluations; collect, translate, and verify evaluation datasets; perform data preparation, analysis, modelling, coding, testing, validation, and deployment; collaborate with cross-functional teams; maintain code repository and documentation standards
Seniority
Mid-level, hands-on IC