AI search news ·
ChatGPT has four hidden retrieval pipelines, and one prompt in nine switches between them — cited sites change with it
In a study published July 5, 2026, Chris Green ran 1,000 prompts ten times each — 9,946 completed runs — and found that ChatGPT's web answers are served by four different internal retrieval sources, none of which OpenAI documents and none of which appear on the citation card you see. 11.6% of prompts switched primary source between otherwise identical runs. When the source switched, the overlap in cited URLs fell by roughly 45%. Same question, same day, same model — a materially different set of businesses cited. If you measure AI visibility for a client, this is the study that explains the noise you have been arguing with.
Primary source: Chris Green (independent study, 9,946 runs)
What the study found
The primary source is Green's write-up, “What can we learn from ChatGPT search?” (July 5, 2026). The method: 1,000 prompts split by intent, each executed 10 times, 9,946 runs completing. Behind the answers sit four labelled search sources, in this distribution:
| Retrieval source | Share of primary search sources | What it appears to specialise in |
|---|---|---|
| Labrador | 88.1% | Smaller, focused retrieval sets; evergreen and educational content |
| Bright | 9.9% | More URLs, more citations, more diverse domains; triggered by current events and date references |
| Oxylabs | 1.7% | Appears infrequently; diverse domain selection |
| SERP | 0.3% | Rare; strongly associated with news queries |
The finding that matters is not the split — it is the instability. 88.4% of prompts kept the same source across all ten runs. The other 11.6% changed primary search source between identical repeated executions, and 13.4% showed changes in their source combinations. When the source changed, the citation set changed with it: URL overlap dropped from 0.273 to 0.149 (about 45%) and domain overlap from 0.265 to 0.155 (about 42%). A second, independent researcher — Suganthan Mohanadasan, working from network traffic rather than scraped sessions — reached the same structural conclusion, which is why we are publishing this rather than holding it.
The caveats, stated plainly
Green is explicit about the limits and we will not launder them: the data comes from Bright-Data-mediated sessions, query fan-out is only partly visible, and OpenAI has never documented what these source labels do — the names are observed, not confirmed by OpenAI. Mohanadasan likewise cautions that his percentages come from a small, tech-skewed batch. So treat the exact percentages as directional. The structural claim — that more than one retrieval pipeline exists, that routing is not deterministic, and that citations move when routing moves — is what two independent methods agree on, and it is the part your measurement has to survive.
What it means for anyone measuring AI visibility
- A single check of a client's AI visibility is not a measurement — it is one sample from a distribution. If roughly one prompt in nine re-routes between runs and takes ~45% of the cited URLs with it, then running a query once and screenshotting the answer tells you almost nothing about whether your client is reliably recommended. Anyone selling AI visibility on the back of a single screenshot is selling a coin flip. This is the same effect we measured from the outside in our own repeat-run data — this study supplies the mechanism.
- Freshness signals decide which pipeline you are even competing in. Bright — the pipeline that pulls more URLs from more domains — is the one that activates on current events and date references. Labrador, which carries 88.1% of the load, favours evergreen, established, educational content. Those are two different content strategies, and which one a client's category lands in is not up to the client.
- Report ranges and rates, not verdicts. The defensible deliverable for a client is “you were named in 4 of 12 answers across repeated runs, here is the trend,” not “ChatGPT recommends you.” The second sentence is not something the system can support — and now there is a 9,946-run study you can hand a client who asks why last month's answer looks different this month.
The takeaway
The engines are not a ranking you can check. They are a distribution you have to sample — repeatedly, across engines, with the same panel of questions every time, or the number you report is noise wearing a suit. That is exactly why our own audit runs every question three times against each of four engines and reports a rate rather than an answer, and why we publish our own zero each Monday from the same panel.
You can run the free 60-second check on any business and see the rate across ChatGPT and Perplexity (the $79 audit adds Gemini and Claude, 25 questions sampled three times each). If you do this for clients, the $249 agency 5-pack runs the full repeated-sample audit across five of them, white-labeled — the version of this you can actually put in front of a client without hedging.
See your number
A free 60-second check shows what AI says about you.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.