FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Friday, 18 September 2026 · Oslo · London · New York

Research note

Six disclosures that make a benchmark result checkable

What is the minimum a published AI result must disclose before a reader can verify it?

Dr. Annika Holm, PhD · Edited by Ingrid Sørensen · Published 2026-09-15 · Last reviewed 2026-09-16

Evidence base

Published evaluation frameworks and reproducibility guidance from standards bodies and academic venues.

A benchmark number without methodology is a claim. Reproducing one requires six things, and the absence of any single one makes the result unverifiable in practice.

First, the dataset and its version, because benchmark datasets are revised. Second, the exact model version, because product names persist across silent updates. Third, the prompts or prompt template, because prompt formatting alone moves scores measurably.

Fourth, the number of runs, because a single run reports a sample and not a performance level. Fifth, the scoring rule, in enough detail that two readers would score the same output the same way. Sixth, exclusions, because a result computed after dropping failed runs is a different result.

These are not demanding requirements. They are the disclosures that academic venues have expected for years and that a substantial share of commercial AI claims still omit. Our forthcoming AI Evaluation Transparency Index scores published vendor claims against exactly these six, and scores disclosure rather than performance.

The distinction is worth restating. A vendor reporting a modest result with complete methodology is more useful to a buyer than a vendor reporting a strong result with none, because only the first can be checked against the buyer's own workload.

What we cannot conclude

  • That a fully disclosed evaluation is a well-designed one. Disclosure permits scrutiny; it does not guarantee quality.
  • That six disclosures are sufficient for every result type; safety and fairness evaluations require more.

This note feeds into AI Evaluation Transparency Index.