FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Thursday, 17 September 2026 · Oslo · London · New York

Research · Long-form Report

Machine Learning Still Struggles to Repeat Its Own Successes

Five years after reproducibility initiatives spread through computer science conferences, verifying claimed results remains expensive, irregular, and rarely rewarded.

A black and white photograph of an empty server rack in a dim data centre with loose cables resting on the concrete floor.
A black and white photograph of an empty server rack in a dim data centre with loose cables resting on the concrete floor.

Independent coverage

A

By Andy Oram

Contributing Writer — Health AI / Open Source · Freelance

Edited by Dr. Annika Holm, PhD

Published 8 September 2026

7 min read

Evidence: Peer-reviewed evidence

Applied machine learning has long promised empirical clarity. Most modern papers present tables filled with bold numbers indicating benchmark improvements. Yet independent teams still find it difficult to reproduce those gains under identical conditions.

The reasons for non-replication are mundane. Unreported random seeds, undisclosed preprocessing scripts, and undocumented hardware configurations alter performance metrics. When code is missing, rebuilding the architecture from text alone rarely yields the published outcome.

The burden of shared artifacts

Major conferences now require code submission checklists and reproducibility declarations. These measures have increased the volume of code repositories linked in preprints. However, repository availability does not guarantee working software.

Many public archives contain incomplete dependencies or point to private internal datasets. Even when dependencies resolve, silent differences in library versions shift numerical precision. A model trained on one accelerator can diverge when run on another.

The compute barrier in verification

Scale introduces a more severe obstacle. Verifying modern foundation models requires millions of dollars in compute infrastructure. Independent university labs cannot afford to run validation sweeps across multiple seeds.

The most influential claims escape third-party verification because the cost of checking them is prohibitive. As a consequence, replication work concentrates on lightweight baselines. Scepticism grows, but auditing remains out of reach for most researchers.

Shifting incentives for applied work

Professional incentives still favor novel claims over verified baselines. Academic tenure committees and corporate hiring teams reward new architecture proposals rather than systematic replications of existing models.

Some labs have started publishing negative results and sensitivity reports. These studies show that fine-tuning gains often disappear when baselines receive equal tuning budgets.

The discipline is maturing, but slowly. Until verifying prior work carries the same prestige as publishing a new benchmark record, the empirical foundation of applied systems will remain fragile.

"The most influential claims escape third-party verification because the cost of checking them is prohibitive."

Published 8 September 2026