FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Thursday, 17 September 2026 · Oslo · London · New York

Research · Long-form Report

Machine learning research is easier to reproduce than psychology and harder than it should be.

Code release rates have improved sharply. Compute cost, undocumented hyperparameters and closed evaluation data still block independent verification of the most consequential claims.

Glassware and a notebook on a laboratory bench
Glassware and a notebook on a laboratory bench

Independent coverage

A

By Andy Oram

Contributing Writer — Health AI / Open Source · Freelance

Edited by Dr. Annika Holm, PhD

Published 8 September 2026

10 min read

Evidence: Analysis

The machine learning field deserves some credit. Code release alongside publication is now normal at the major conferences, artefact review tracks exist, and the community norm has shifted decisively within a decade.

The remaining obstacles are structural rather than cultural, and they concentrate precisely on the claims that matter most.

Cost as a barrier to verification

A frontier scale training run cannot be independently repeated by a university group. Verification therefore relies on the reported numbers, on downstream behaviour, and on trust. That is an uncomfortable position for a field that describes itself as empirical.

The hyperparameter problem

Reproduction failures in mid-scale work are dominated by details that were never written down. Learning rate schedules adjusted by hand, data filtering steps applied before the documented pipeline, and random seeds selected after the fact.

None of this is fraud. It is the ordinary residue of research conducted under deadline, and it is exactly what a well specified artefact submission catches.

Evaluation data that cannot be inspected

Where the evaluation set is private, an outside reader cannot assess contamination, difficulty or representativeness. Private sets are defensible as an anti-gaming measure. They are also unfalsifiable, and the field has not settled how to weigh that trade.

What improves matters concretely

Three things, all achievable. Mandatory reporting of compute and wall clock alongside results, so cost of verification is visible. Seed and schedule disclosure as a condition of acceptance. And a standing fund for replication work, which currently has almost no career value and therefore almost no volunteers.

Why it matters outside academia

Because procurement, regulation and clinical deployment all cite this literature. A claim that cannot be checked is still a claim that ends up in a tender document, and eventually in a decision about a patient or a grid.

"A result that costs two million dollars to check is not a result. It is an announcement with citations."

Sources

Published 8 September 2026