Machine Learning Engineer, Evaluation
Core
Designing and building LLM-powered evaluation pipelines to measure developer skills in an AI-assisted era, ensuring consistency, fairness, and scalability across hundreds of thousands of assessments.
Role type
Senior IC machine-learning engineer (evaluation methodology & LLM systems)
Builds
LLM-powered evaluation pipelines, RAG pipelines, fine-tuning workflows, and benchmarking infrastructure for skill assessment
Domain
EdTech / AI Evaluation / Generative AI
Deliverable
production ML models
Required skills
LLM system design, RAG pipeline construction, model fine-tuning, experimental design, bias detection and auditing, rubric definition, system architecture, data pipeline engineering, model monitoring, plain-language technical translation
Preferred skills
Generative AI evaluation framework experience, psychometrics or educational assessment background, LLM benchmarking/alignment research, research-to-product shipping
Technologies
LLMs, RAG, fine-tuning frameworks, benchmarking tools
Responsibilities
Build LLM-powered evaluation pipelines for consistent AI usage skill assessment; Own end-to-end evaluation methodology including rubrics, model application, and bias audits; Design and run experiments to define effective evaluation standards; Build RAG pipelines and fine-tuning workflows for reliable model adherence; Define benchmarking infrastructure to track evaluation quality and catch regressions; Translate model behavior into understandable outcomes for product managers and candidates
Seniority
Senior, hands-on IC with research mindset