Skill tests for AI data work
How many data annotators would pass a real test?
Most teams find out the hard way, after the bad labels are already in the training set. mindsparx tests your reviewers on the real job and grades every answer against an expert answer key, so you find out first.
Run a free pilot this month
Get a real competency report on your annotation team, free. Leave your email and I'll reply personally to get started. No meeting, no contract.
Data annotation, LLM output evaluation, transcription, RLHF preference comparison: the human-in-the-loop work that decides the quality of AI training data, long before it reaches model training or model evaluation. Data points no resume can prove.
For data labeling vendors, RLHF and human-in-the-loop providers, and AI lab evaluation teams.
Nobody really checks this work.
Resumes don't prove skill.
Three years of annotation experience on paper says nothing about whether someone catches a subtle factual error or applies a labeling rubric the same way twice. It's a claim, not a measurement.
Vendors grade their own homework.
In most data labeling teams, quality assurance means the company that supplied the annotators also checks the annotators. That is not an independent check, and everyone in the industry knows it.
Problems surface too late.
Bad training data gets noticed after the model is trained or the client complains. By then you still don't know which reviewer made the mistakes, or what kind they were.
What you get
mindsparx tests them with real tasks, graded automatically against expert answer keys. You get a report with data points on industry-standard metrics and a clear score per person. And you can create your own work types that map to the same metrics.
Tests shaped like the job
Judge a model response, annotate real data, transcribe real audio, compare two outputs. No multiple choice, no trivia.
Graded against expert answers
Every submission is scored automatically against a validated answer key. There is no grader in the room to charm.
Reports on industry-standard metrics
Accuracy against the key, word error rate, problems caught. Data points the industry already trusts, and a clear score per person.
Anonymous by design
Every candidate is a code, never a name. You keep the list of who is who. We never see it.
Your work types, your rules
Run a mindsparx standard as it is, copy one and change everything, or build your own work type that maps to the same industry-standard metrics.
Certificates anyone can verify
Whoever passes gets a public certificate with the score, the pass bar and the date. One click to check it, no login, no PDF to fake.
How it works
- 1
A test gets built.
Send your annotation guidelines or pick a work type: AI output evaluation, data annotation, transcription or preference comparison. AI drafts the questions from your rules, a human validates every answer in the key.
- 2
People take it through a link.
One link per person, identified only by a code. No signup, no name, no email. A test takes about half an hour.
- 3
Scores come back, passers get proof.
Each person is scored on accuracy, completeness, consistency and more. Anyone who clears the pass bar earns a verifiable certificate.
The testing engine
How the grading works.
Every test is published with an answer key a human expert validated. When a candidate submits, the engine grades on the server with the metric that fits the work type: accuracy against the key for annotation, word error rate for transcription, problems caught for output evaluation.
Scores roll up into competencies and one composite per person, judged against a fixed pass bar that is the same for everyone. The answer key never leaves the server, and nobody grades their own work.
grading submission · code TR-014
work type: data annotation
items answered 24 / 24
accuracy vs key 92%
problems caught (F1) 0.86
word error rate 4.1%
composite score 88
pass bar 75
→ PASS · certificate issued
mindsparx.ai/cert/8ccb27…Common questions
What is mindsparx?
mindsparx is a skill testing platform for AI data work, also called human-in-the-loop or AI training data work. It tests data annotation, LLM output evaluation, transcription and RLHF preference comparison with real tasks, grades every answer against an expert answer key, and issues verifiable certificates to the people who pass.
How are the tests graded?
Automatically, against an answer key a human expert has validated. Scoring happens on the server, candidates never see the answer key, and the pass bar is an absolute score, not a curve.
What is a verifiable certificate?
A public page at its own unguessable address showing the assessment, the candidate's code, the score, the pass bar and the date. Anyone can open the link and confirm the certificate is genuine, with no account and no login.
How do you test annotators without collecting their personal data?
Every candidate is an anonymous code. The company running the test keeps the list of which code belongs to which person. mindsparx never sees names or emails.
What does a pilot cost?
Nothing. Send 5 to 10 codes from your team and you get back a real competency report, free, with no contract.
Can I take a test as an individual?
Not yet on your own. Today assessments are run by the companies you work with. Individual sign-up is coming, and the email list gets first access.