AI QA Trainer - Freelance Project
Core
Evaluating large-scale language models to ensure factual accuracy, logical soundness, safety, and reliability across real-world scenarios.
Role type
AI QA Trainer (LLM Evaluation & Safety)
Builds
Enterprise-grade AI platforms and evaluation frameworks
Domain
Artificial Intelligence / Large Language Models (LLMs)
Deliverable
production ML models
Required skills
LLM evaluation, adversarial testing/red-teaming, regression testing at scale, bias/fairness auditing, grounding verification, prompt engineering, test automation (Python/SQL), high-signal bug reporting
Preferred skills
Evaluation rubric design, chain-of-reasoning reliability analysis, tool-use correctness verification, retrieval-augmentation fidelity checks
Technologies
Python, SQL, PyTest, OpenAI Evals, W&B
Responsibilities
Converse with models on real-world scenarios to verify factual accuracy and logical soundness; design and run test plans and regression suites; build clear rubrics and pass/fail criteria; capture reproducible error traces with root-cause hypotheses; partner on adversarial red-teaming and automation; dashboard quality deltas over time
Seniority
Mid-Senior Level, hands-on IC