FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Thursday, 17 September 2026 · Oslo · London · New York

AI · Analysis

The benchmark industry is failing, and the people who run it say so first.

Contamination, saturation and selective reporting have made public leaderboards a weak guide to production behaviour. Evaluation researchers are rebuilding around task-specific suites.

Laboratory glassware and a notebook on a bench
Laboratory glassware and a notebook on a bench

Independent coverage

M

By Martin Anderson

Contributing Writer — AI / ML · Freelance

Edited by Dr. Annika Holm, PhD

Published 14 September 2026

8 min read

Evidence: Analysis

The most sceptical people about public model benchmarks are the researchers who build them. That is not a rhetorical flourish. Read the limitations sections of the major evaluation papers and you find an unusually frank body of work.

Three problems recur. Contamination, where test items leak into training data. Saturation, where scores cluster so tightly at the top that differences are noise. And selective reporting, where a vendor publishes the seven suites it leads on.

Why this matters outside research

Procurement teams use leaderboards. A public score becomes a line in a tender document, and a difference of two points becomes a justification for a contract worth millions. The number was never built to carry that weight.

What replaced it in serious shops

Task-specific evaluation sets built from the organisation's own historical work, held privately, refreshed quarterly. A hundred real examples from your own domain, graded by people who know the domain, beats any public suite for procurement purposes.

It is more work. It is also the only approach that survives a model update, because you can rerun it the day the vendor ships.

A reasonable position

Treat public benchmarks as a screening filter and nothing more. They are good at telling you which models are not worth testing. They are poor at telling you which one to buy.

"A benchmark that appears in the training corpus measures memory, not capability, and nobody can prove it did not."

Sources

Published 14 September 2026