Member of Technical Staff, Data Analysis and Evaluation
Core
Designing data collection tasks, evaluating dataset quality, and analyzing the robustness and generalisability of large language models (LLMs) to ensure high-quality training data and reliable model performance.
Role type
Senior IC data analysis and evaluation engineer (LLMs)
Builds
High-quality datasets and robust, scalable LLM systems for developers and enterprises
Domain
Artificial Intelligence / Large Language Models / Data Science
Deliverable
production ML models
Required skills
Statistical methods, experimental design, data collection task design, human annotator management, dataset quality assessment, LLM training and fine-tuning, distributed training infrastructure, Python programming, ML frameworks (PyTorch, TensorFlow, JAX), generalisability and robustness analysis
Preferred skills
Publications at top-tier venues (NeurIPS, ICML, ICLR, etc.)
Technologies
Python, PyTorch, TensorFlow, JAX
Responsibilities
Design and oversee data collection tasks including supporting human annotators; Develop and apply statistical methods to evaluate dataset quality; Analyse and assess the generalisability and robustness of ML systems; Train and fine-tune LLMs on distributed training infrastructures; Conduct experiments to evaluate model performance and identify areas for improvement
Seniority
Senior, hands-on IC