AI QA Trainer - LLM Evaluation - Freelance Project
Core
Evaluating large-scale language models for enterprise-grade reliability, safety, and reasoning accuracy through rigorous testing and red-teaming.
Role type
Senior AI QA Trainer (LLM Evaluation)
Builds
Evaluation data, test plans, rubrics, and quality metrics for LLMs serving enterprise workflows.
Domain
Artificial Intelligence / Large Language Models / Model Safety
Deliverable
production ML models
Required skills
LLM evaluation, adversarial red-teaming, prompt engineering, bias/fairness auditing, grounding verification, test automation (Python/SQL), regression testing, bug reporting
Preferred skills
Shipped QA for ML/AI systems, experience with LLM eval tooling (OpenAI Evals, RAG evaluators, W&B), chain-of-reasoning analysis
Technologies
Python, SQL, OpenAI Evals, RAG evaluators, W&B
Responsibilities
Converse with models on real-world scenarios to verify factual accuracy and logical soundness; design and run test plans and regression suites; build rubrics and pass/fail criteria; capture reproducible error traces with root-cause hypotheses; partner on adversarial red-teaming and automation; dashboard quality deltas over time
Seniority
Mid-Senior Level, hands-on IC