Research note
Six disclosures that make a benchmark result checkable
What is the minimum a published AI result must disclose before a reader can verify it?
Dr. Annika Holm, PhD · Edited by Ingrid Sørensen · Published 2026-09-15 · Last reviewed 2026-09-16
Evidence base
Published evaluation frameworks and reproducibility guidance from standards bodies and academic venues.
A benchmark number without methodology is a claim. Reproducing one requires six things, and the absence of any single one makes the result unverifiable in practice.
First, the dataset and its version, because benchmark datasets are revised. Second, the exact model version, because product names persist across silent updates. Third, the prompts or prompt template, because prompt formatting alone moves scores measurably.
Fourth, the number of runs, because a single run reports a sample and not a performance level. Fifth, the scoring rule, in enough detail that two readers would score the same output the same way. Sixth, exclusions, because a result computed after dropping failed runs is a different result.
These are not demanding requirements. They are the disclosures that academic venues have expected for years and that a substantial share of commercial AI claims still omit. Our forthcoming AI Evaluation Transparency Index scores published vendor claims against exactly these six, and scores disclosure rather than performance.
The distinction is worth restating. A vendor reporting a modest result with complete methodology is more useful to a buyer than a vendor reporting a strong result with none, because only the first can be checked against the buyer's own workload.
What we cannot conclude
- That a fully disclosed evaluation is a well-designed one. Disclosure permits scrutiny; it does not guarantee quality.
- That six disclosures are sufficient for every result type; safety and fairness evaluations require more.
This note feeds into AI Evaluation Transparency Index.