AI · Analysis
The benchmark industry is failing, and the people who run it say so first.
Contamination, saturation and selective reporting have made public leaderboards a weak guide to production behaviour. Evaluation researchers are rebuilding around task-specific suites.

Independent coverage
Published 14 September 2026
8 min read
Evidence: Analysis
The most sceptical people about public model benchmarks are the researchers who build them. That is not a rhetorical flourish. Read the limitations sections of the major evaluation papers and you find an unusually frank body of work.
Three problems recur. Contamination, where test items leak into training data. Saturation, where scores cluster so tightly at the top that differences are noise. And selective reporting, where a vendor publishes the seven suites it leads on.
Why this matters outside research
Procurement teams use leaderboards. A public score becomes a line in a tender document, and a difference of two points becomes a justification for a contract worth millions. The number was never built to carry that weight.
What replaced it in serious shops
Task-specific evaluation sets built from the organisation's own historical work, held privately, refreshed quarterly. A hundred real examples from your own domain, graded by people who know the domain, beats any public suite for procurement purposes.
It is more work. It is also the only approach that survives a model update, because you can rerun it the day the vendor ships.
A reasonable position
Treat public benchmarks as a screening filter and nothing more. They are good at telling you which models are not worth testing. They are poor at telling you which one to buy.
"A benchmark that appears in the training corpus measures memory, not capability, and nobody can prove it did not."
Sources