2 Machine Learning Engineer Evaluation
Core
Building LLM-powered evaluation pipelines and infrastructure to assess developer skills in an AI-assisted environment, ensuring consistency, fairness, and scalability across hundreds of thousands of assessments.
Role type
Senior Machine Learning Engineer (Evaluation & Benchmarking)
Builds
LLM-powered evaluation pipelines, RAG pipelines, fine-tuning workflows, and benchmarking infrastructure for skill assessment.
Domain
AI/ML evaluation, software engineering assessment, generative AI
Deliverable
production ML models
Required skills
LLM system design, RAG pipeline development, model fine-tuning, experimental design, bias detection and auditing, system architecture, data pipeline engineering, model monitoring, rubric design, stakeholder communication
Preferred skills
Generative AI evaluation frameworks, psychometrics, educational assessment, human-in-the-loop evaluation, LLM benchmarking, research-to-product translation
Technologies
LLMs, RAG, fine-tuning frameworks, benchmarking tools
Responsibilities
Build LLM-powered evaluation pipelines, own the evaluation methodology end-to-end, design and run experiments to define evaluation standards, build RAG pipelines and fine-tuning workflows, define benchmarking infrastructure, translate model behavior into understandable outcomes
Seniority
Senior, hands-on IC with research mindset
