Machine Learning Engineer – ML Evaluation & Experiment Design
Core
Reviewing and evaluating machine learning challenges, datasets, and pipelines to ensure they are technically sound, reproducible, and require genuine ML reasoning.
Role type
Senior IC machine learning engineer (evaluation & experiment design)
Builds
Rigorous ML challenges, benchmark datasets, and evaluation pipelines for AI model training
Domain
Applied machine learning, data-centric AI, experiment design
Deliverable
production ML models
Required skills
ML experiment design, model selection, hyperparameter tuning, model evaluation, data preprocessing and validation, train/validation/test split methodology, data leakage detection, label noise identification, distribution shift analysis, statistical significance testing, debugging ML workloads across CPU/GPU
Preferred skills
Kaggle/DrivenData competition experience, benchmark dataset design, synthetic data generation, statistical testing (confidence intervals/effect sizes), RLHF/AI model evaluation, ML curriculum development, understanding of ML failure modes (shortcut learning, Goodhart's Law, Simpson's paradox)
Technologies
N/A
Responsibilities
Reviewing ML challenges for technical soundness and solvability, evaluating datasets for meaningful signals and artifacts, detecting metric gaming and evaluation flaws, verifying reproducibility across data-to-evaluation pipelines, assessing challenge difficulty calibration, providing recommendations for task improvement or exclusion
Seniority
Senior, hands-on IC