Data Annotation Specialist, Generalist
Core
Evaluate, rank, and stress-test Large Language Model outputs to improve model performance and safety across text, image, and structured data.
Role type
Independent contractor data annotation specialist (LLM evaluation)
Builds
High-quality training and evaluation datasets for enterprise AI models
Domain
Artificial Intelligence / Large Language Models
Deliverable
production ML models
Required skills
Judgment-driven evaluation, adversarial probing, rubric design, multimodal data annotation, consistency calibration, experimental task adaptation, performance reporting
Preferred skills
Experience with LLM failure modes (hallucination, sycophancy), fluency in a second language
Technologies
JSON, CSV, TSV, Markdown, XML, YAML
Responsibilities
Evaluate and rank model outputs based on accuracy, helpfulness, tone, and safety; stress-test models to surface failure modes and unsafe behavior; create datasets by authoring prompts and responses; build and apply rubrics and taxonomies for grading; annotate and correct multimodal data; calibrate standards through inter-annotator agreement checks; report on model performance trends
Seniority
Entry-level to Mid-level, hands-on IC
