What we test

Skill tests for every kind of AI data work

Every test below is a real, job-shaped task, not a quiz. Answers are graded automatically against an expert-validated answer key, using the metric the industry uses for that task. Candidates stay anonymous codes, and everyone who clears the pass bar earns a verifiable certificate. Start from any of these as-is, or adapt one to your own guidelines.

AI Response Evaluation

Can your reviewers judge LLM outputs reliably? They score real model answers against your guidelines and flag the problems that are actually there.

Judge the quality of AI answers

LLM output evaluation (judge + flag)

overall score vs gold band · problems-caught F1

Compare & Rank AI Answers

The core RLHF skills: pick the better of two model responses, or put several in order, judged against expert preferences.

Pick the better of two AI answers

Pairwise preference (RLHF)

agreement with expert preference

Rank several AI answers best to worst

Response ranking (RLHF)

rank agreement vs expert order (Spearman footrule)

Search & Relevance

Query understanding and result grading, the way search relevance programs actually run.

Classify what a search query is looking for

Query intent classification

exact match vs gold

Grade how relevant a result is to a query

Graded relevance (SERP rating)

exact match vs gold grade

Content Moderation & Safety

Harm taxonomy decisions, severity calls, and finding personal information that must be redacted.

Classify harmful content and its severity

Harm taxonomy classification & severity grading

set overlap (categories) · exact match (severity)

Mark personal information in text

PII redaction

span F1 vs gold spans

Text Labeling

Classic annotation: classification and marking entities in text, the ground-truth work behind most training data.

Put text into the right category

Text / intent classification

exact match vs gold

Mark the entities in a text

NER / span tagging

span F1 vs gold spans

Translation & Localization QA

Translation review the industry way: MQM error types and severity, and machine-translation post-editing.

Judge translation errors and their severity

MQM error classification & severity

exact match (severity) · set overlap (error types)

Fix a machine translation with minimal edits

Post-editing (MTPE)

word similarity to gold post-edit (HTER-style)

Transcription & Audio

Type what is said, tag what is heard. Scored by the standard speech metrics.

Type out exactly what is said in audio

Verbatim transcription

word error rate (WER)

Tag the sounds or events in an audio clip

Audio event tagging

set overlap vs gold tags

Fact & Data Verification

Is the extracted value right? Is the AI answer actually supported by the source? Increasingly the highest-volume human-in-the-loop task.

Check whether extracted data is actually correct

Structured-attribute verification

verdict exact match · hallucination recall

Is this answer supported by the source?

Groundedness / citation faithfulness (RAG)

label agreement vs gold

Image & Video

Image classification and moderation with the same grading engine. Box and polygon drawing is available on request.

Classify or moderate images

Image classification / attribute tagging

exact match · set overlap

Writing Tasks

Prompt and response writing for supervised fine-tuning, collected under test conditions for expert review.

Write an ideal response to a prompt

Demonstration / SFT authoring

collected for expert review (not auto-scored)

Available on request

Some work types need specialized annotation screens. We build these against a live client project rather than in the abstract:

  • · Bounding boxes, polygons and keypoints on images (IoU scored)
  • · Speaker diarization and audio timestamping
  • · Video object tracking

See it on your own team, free

Pick a work type, send 5 to 10 candidate codes, and get a full competency report back. No meeting, no contract.