What we test
Skill tests for every kind of AI data work
Every test below is a real, job-shaped task, not a quiz. Answers are graded automatically against an expert-validated answer key, using the metric the industry uses for that task. Candidates stay anonymous codes, and everyone who clears the pass bar earns a verifiable certificate. Start from any of these as-is, or adapt one to your own guidelines.
AI Response Evaluation
Can your reviewers judge LLM outputs reliably? They score real model answers against your guidelines and flag the problems that are actually there.
Judge the quality of AI answers
LLM output evaluation (judge + flag)
overall score vs gold band · problems-caught F1
Compare & Rank AI Answers
The core RLHF skills: pick the better of two model responses, or put several in order, judged against expert preferences.
Pick the better of two AI answers
Pairwise preference (RLHF)
agreement with expert preference
Rank several AI answers best to worst
Response ranking (RLHF)
rank agreement vs expert order (Spearman footrule)
Search & Relevance
Query understanding and result grading, the way search relevance programs actually run.
Classify what a search query is looking for
Query intent classification
exact match vs gold
Grade how relevant a result is to a query
Graded relevance (SERP rating)
exact match vs gold grade
Content Moderation & Safety
Harm taxonomy decisions, severity calls, and finding personal information that must be redacted.
Classify harmful content and its severity
Harm taxonomy classification & severity grading
set overlap (categories) · exact match (severity)
Mark personal information in text
PII redaction
span F1 vs gold spans
Text Labeling
Classic annotation: classification and marking entities in text, the ground-truth work behind most training data.
Put text into the right category
Text / intent classification
exact match vs gold
Mark the entities in a text
NER / span tagging
span F1 vs gold spans
Translation & Localization QA
Translation review the industry way: MQM error types and severity, and machine-translation post-editing.
Judge translation errors and their severity
MQM error classification & severity
exact match (severity) · set overlap (error types)
Fix a machine translation with minimal edits
Post-editing (MTPE)
word similarity to gold post-edit (HTER-style)
Transcription & Audio
Type what is said, tag what is heard. Scored by the standard speech metrics.
Type out exactly what is said in audio
Verbatim transcription
word error rate (WER)
Tag the sounds or events in an audio clip
Audio event tagging
set overlap vs gold tags
Fact & Data Verification
Is the extracted value right? Is the AI answer actually supported by the source? Increasingly the highest-volume human-in-the-loop task.
Check whether extracted data is actually correct
Structured-attribute verification
verdict exact match · hallucination recall
Is this answer supported by the source?
Groundedness / citation faithfulness (RAG)
label agreement vs gold
Image & Video
Image classification and moderation with the same grading engine. Box and polygon drawing is available on request.
Classify or moderate images
Image classification / attribute tagging
exact match · set overlap
Writing Tasks
Prompt and response writing for supervised fine-tuning, collected under test conditions for expert review.
Write an ideal response to a prompt
Demonstration / SFT authoring
collected for expert review (not auto-scored)
Available on request
Some work types need specialized annotation screens. We build these against a live client project rather than in the abstract:
- · Bounding boxes, polygons and keypoints on images (IoU scored)
- · Speaker diarization and audio timestamping
- · Video object tracking
See it on your own team, free
Pick a work type, send 5 to 10 candidate codes, and get a full competency report back. No meeting, no contract.