Sample report
This is the report you receive after your candidates sit an assessment: the whole run ranked, any code opened into a full breakdown, and a verifiable certificate for everyone who clears the bar. Shown here for an AI output evaluation test with 6 reviewers and a pass bar of 70%. Every number and code on this page is fictional; a real report uses your team's own codes and the work type you choose, and reviewers are always a code, never a name or an email.
Independent candidates who take a standard on their own account see a lighter version: the verdict, the competency profile and the same did-well and needs-improvement patterns, without the question-by-question table. That table goes only to the company that runs the test with them.
Section 1
AI Output Evaluation (English) · Form A. A run always serves exactly one test, so every row below sat the same questions. 8 candidates were invited and 6 finished; the two who have not yet finished sit in a separate list with their invite links so you can chase them.
6
4/6
68.4
28m
2
Candidate | Result | Accuracy | Consistency | Completeness | Efficiency | Calibration | Time |
|---|---|---|---|---|---|---|---|
| TR-014 | Pass 91% | 91% | 90% | 84% | 76% | 82% | 34m |
| TR-007 | Pass 84% | 84% | 88% | 79% | 81% | 76% | 41m |
| TR-022 | Pass 77% | 77% | 74% | 71% | 68% | 70% | 38m |
| TR-009 | Pass 72% | 72% | 70% | 61% | 74% | 63% | 29m |
| TR-018 | Fail 58% | 58% | 66% | 42% | 58% | 51% | 22m |
| TR-011 | Fail 46% | 46% | 38% | 30% | 95% | 34% | 6m |
Each row opens into a full per-candidate report; the integrity flags raised on a sitting and the candidate's ranking within this run are carried there.
Section 2
Every code in the table opens into this. Estimated duration of this test: ~40 min.
84% agreement with the validated answer key. The bar for this assessment is 70%.
Ranked #2 of 6 in this run
Caught every factual error in the test, 4 of 4.
Question 5: flagged the wrong launch year in an AI answer that read fluently.
Held the same standard from the first item to the last.
Consistency 88, their strongest competency.
Sometimes flags problems that are not there, 2 times in 12 items.
Question 9: marked the answer off-topic when it was on-topic.
Confidence ran ahead of correctness on the hardest items.
Calibration 76, their lowest competency.
| # | Question | Score → expected | Problems flagged | Verdict |
|---|---|---|---|---|
| 2 | Summary of a news article | 4→4 | Unsupported claim | ● |
| 5 | Fluent answer about a product launch | 2→2 | Factual errorIgnores an instruction | ◐ |
| 9 | Long answer to a broad question | 3→4 | Off-topic | ◐ |
Section 3
Everyone who clears the bar gets a certificate with its own public verification link. The page shows the candidate code, the assessment, the score and the pass bar, and a status banner that reads VALID, REVOKED or EXPIRED. Anyone you forward the link to can check it in one click; no account needed.
Certificates are alive, not screenshots. If the same candidate later sits a retest and lands below the bar, the certificate is revoked automatically and the public page says so immediately. And because the whole platform runs on codes, the certificate vouches for a code and a score; connecting that code to a person stays in your hands.
Browse the test catalog to see the kinds of work these runs can cover.
A pilot is one assessment built on your own guidelines, run against your candidate codes. You get this report back, ranked, with certificates for whoever clears the bar.
Request a pilot