Research Engineer I (AI Evaluation, Trust Technologies)
Core
Design evaluation frameworks, test datasets, and scoring methodologies for AI/ML systems; conduct error analysis and build supporting prototypes to ensure AI systems are trustworthy, fair, and safe in real-world conditions.
Role type
Research Engineer (AI Evaluation & Validation)
Builds
Evaluation frameworks, test datasets, annotation guidelines, and tooling for AI/ML systems
Domain
AI Safety, Trust Technologies, Large Language Models (NLP) (via careerplan.io/jobs/R00025937-research-engineer-i-ai-evaluation-trust-technologies-at-ntu)
Deliverable
production ML models
Required skills
Python, ML workflows, model evaluation, error analysis, NLP, large language models, data annotation, experiment design
Preferred skills
Evaluation methodology, data annotation
Technologies
Python, Large Language Models
Responsibilities
Design evaluation frameworks and test datasets; develop annotation guidelines and monitor labelling quality; conduct error analysis and translate findings into recommendations; run structured experiments with large language models; build supporting prototypes and tooling; engage with external partners and clients on requirements and technical meetings.
Seniority
Individual Contributor, Mid-Level
