FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Friday, 18 September 2026 · Oslo · London · New York

Research note

A single prompt is an observation, not a ranking

How many repetitions does it take before an AI answer can be described as a stable position?

Nathaniel "Nate" Whitaker · Edited by Dr. Annika Holm, PhD · Published 2026-09-12 · Last reviewed 2026-09-16

Evidence base

Provider documentation on sampling parameters, published work on prompt sensitivity, and the standard statistics of proportion estimates.

Generative systems sample from a probability distribution over tokens. Unless sampling is disabled, two identical requests can return different text, and the difference is frequently material: a named vendor appears in one answer and not the next.

This is not a defect. It follows directly from how decoding works. It does mean that the common commercial practice of querying an assistant once and reporting the result as a position is closer to a single coin flip than to a ranking.

The statistics here are ordinary. Estimating how often an entity appears is estimating a proportion, and the uncertainty in that estimate narrows roughly with the square root of the number of runs. Ten runs leaves a very wide interval. Twenty narrows it usefully. Fifty narrows it further at a cost that is still trivial for a small prompt set.

There is a second source of movement that repetition does not fix. Retrieval-backed systems query a live index, and the index changes. Provider model versions also change behind stable product names. A time series therefore needs the model version and collection date recorded on every run, otherwise a genuine change and a silent provider update are indistinguishable after the fact.

Our working standard for the AI Recommendation Index follows from this: a minimum of twenty runs per prompt per system per collection window, version and locale recorded per run, and frequency reported with the run count attached. We do not publish an ordinal position.

What we cannot conclude

  • That twenty runs is sufficient for every prompt. Prompts with many plausible answers need more.
  • That variation observed through a consumer interface is purely sampling. Personalisation and regional routing cannot be excluded from outside.
  • That frequency of appearance relates to market share, product quality or suitability for any buyer.

This note feeds into AI Recommendation Index.