Vol. XVI · No. 261Friday 18 September 2026World Edition
TheNewsRupt coat of arms crest

The NewsRupt

← Research notes

Ask the same model the same question three times and the lead name changes a third of the time

A short measurement made while building the AI Recommendation Index: repetition, not model choice, is the biggest source of variance in AI vendor recommendations.

Published 18 Sep 2026 · Feeds the AI Recommendation Index
The finding

Across 36 model-prompt cells sampled three times each, 13 (36.1%) returned more than one distinct first-named vendor. Instability differed sharply by model: 17% for gemini-3.8-flash, 42% for gemini-3.1-flash-lite, 50% for gpt-5.4-mini.

Method

Twelve buying questions across three infrastructure categories were sent to three models, three times each, with no history and no tools. For each answer we recorded the first candidate vendor named. A cell — one model, one prompt — counts as unstable if its three runs produced more than one distinct first name.

Reading it

Anyone selling "AI visibility" services runs into this immediately, whether or not they say so. A single query to a single assistant tells you almost nothing about where a brand stands, because the assistant does not have a settled view to report. The variance is not noise around a true ranking; at the top of a concentrated category the lead name barely moves, and in the second tier it moves constantly.

The practical consequence is a sampling floor. If a measurement of brand presence in AI answers is based on fewer than several repetitions per prompt per model, its movement between periods cannot be distinguished from resampling the same distribution. Our own index uses three repetitions, which is enough to demonstrate instability and not enough to estimate it precisely — a limitation we would rather state than hide behind a chart.

It also argues against single-assistant audits. In the same run, Replicate appeared in 75% of one model's inference-provider answers and in none of the other two models' answers. A brand measuring only ChatGPT, or only Gemini, is measuring one retrieval-and-generation pipeline and calling it the market.

Limitations
  • Three repetitions per cell gives a wide confidence interval; 36% should be read as a floor.
  • Models were sampled at their default settings through one gateway; temperature and serving configuration were not controlled.
  • First-named position is a proxy for recommendation strength, not a measure of it.
Sources
Part of a franchise

This note came out of the AI Recommendation Index, where the full dataset and method are published.