AI search news ·

ChatGPT has four hidden retrieval pipelines, and one prompt in nine switches between them — cited sites change with it

In a study published July 5, 2026, Chris Green ran 1,000 prompts ten times each — 9,946 completed runs — and found that ChatGPT's web answers are served by four different internal retrieval sources, none of which OpenAI documents and none of which appear on the citation card you see. 11.6% of prompts switched primary source between otherwise identical runs. When the source switched, the overlap in cited URLs fell by roughly 45%. Same question, same day, same model — a materially different set of businesses cited. If you measure AI visibility for a client, this is the study that explains the noise you have been arguing with.

What the study found

The primary source is Green's write-up, “What can we learn from ChatGPT search?” (July 5, 2026). The method: 1,000 prompts split by intent, each executed 10 times, 9,946 runs completing. Behind the answers sit four labelled search sources, in this distribution:

Retrieval sourceShare of primary search sourcesWhat it appears to specialise in
Labrador88.1%Smaller, focused retrieval sets; evergreen and educational content
Bright9.9%More URLs, more citations, more diverse domains; triggered by current events and date references
Oxylabs1.7%Appears infrequently; diverse domain selection
SERP0.3%Rare; strongly associated with news queries

The finding that matters is not the split — it is the instability. 88.4% of prompts kept the same source across all ten runs. The other 11.6% changed primary search source between identical repeated executions, and 13.4% showed changes in their source combinations. When the source changed, the citation set changed with it: URL overlap dropped from 0.273 to 0.149 (about 45%) and domain overlap from 0.265 to 0.155 (about 42%). A second, independent researcher — Suganthan Mohanadasan, working from network traffic rather than scraped sessions — reached the same structural conclusion, which is why we are publishing this rather than holding it.

The caveats, stated plainly

Green is explicit about the limits and we will not launder them: the data comes from Bright-Data-mediated sessions, query fan-out is only partly visible, and OpenAI has never documented what these source labels do — the names are observed, not confirmed by OpenAI. Mohanadasan likewise cautions that his percentages come from a small, tech-skewed batch. So treat the exact percentages as directional. The structural claim — that more than one retrieval pipeline exists, that routing is not deterministic, and that citations move when routing moves — is what two independent methods agree on, and it is the part your measurement has to survive.

What it means for anyone measuring AI visibility

The takeaway

The engines are not a ranking you can check. They are a distribution you have to sample — repeatedly, across engines, with the same panel of questions every time, or the number you report is noise wearing a suit. That is exactly why our own audit runs every question three times against each of four engines and reports a rate rather than an answer, and why we publish our own zero each Monday from the same panel.

You can run the free 60-second check on any business and see the rate across ChatGPT and Perplexity (the $79 audit adds Gemini and Claude, 25 questions sampled three times each). If you do this for clients, the $249 agency 5-pack runs the full repeated-sample audit across five of them, white-labeled — the version of this you can actually put in front of a client without hedging.

See your number

A free 60-second check shows what AI says about you.

Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.

Run a free AI visibility check

Begin your check

Free · 60 sec

No account · No card · 3 buyer questions, 2 engines

By running a check you agree to our Terms and Privacy Policy.

Who runs this