AI search news ·
Does ChatGPT cite the same sources every time you ask? Only 7.4% of its cited domains survive three identical asks — Perplexity keeps 63.8%
Ask an AI engine the identical question three times in a row and only 25.4% of the domains it cites appear in all three answers — 57.0% appear exactly once. That is the pooled reading from our production instrument: 12 pinned small-business buyer questions × 4 engines × 3 identical samples, run in three separate dated weeks (432 answers, 2,623 cited-domain slots). The per-engine spread is the finding: Perplexity keeps 63.8% of its cited domains across all three asks; ChatGPT keeps 7.4% and Gemini 6.6% — a 9.7× consistency gap. If your AI visibility evidence is one screenshot of one ChatGPT answer, roughly three quarters of what it shows is a one-off draw.
Primary source: AskedAbout — within-run sample-consistency cut of three dated four-engine runs (2026-07-28 · 2026-08-03 · 2026-08-10)
The instrument
Our weekly self-audit asks 12 pinned buyer questions — "What are the best AI visibility tools for small businesses?", "What is a good Profound alternative for a small business?", and ten more — to ChatGPT, Perplexity, Gemini and Claude through their APIs, three identical samples per question per engine, and stores every answer with its returned citation list. That design exists so a citation only counts as stable when it shows up in repeated samples. It also makes the instrument a consistency meter for the engines themselves: within a single run, the three samples of a cell are the same question, same engine, same model, same day, minutes apart. This cut asks one question of that data: of the domains an engine cites at least once across the three asks, what share appears in all three — and what share in only one?
The numbers: a 9.7× consistency gap between engines
Pooled across all three weekly runs (each run: 12 questions × 4 engines × 3 samples = 144 answers; cells where all three samples returned zero citations are excluded and counted separately below):
| Engine | Question-cells measured | Distinct domains cited | Cited in all 3 asks | Cited in only 1 ask | Mean pairwise overlap (Jaccard) |
|---|---|---|---|---|---|
| Perplexity | 36 of 36 | 751 | 479 (63.8%) | 164 (21.8%) | 0.758 |
| Claude | 28 of 36 | 343 | 82 (23.9%) | 177 (51.6%) | 0.417 |
| ChatGPT | 36 of 36 | 544 | 40 (7.4%) | 412 (75.7%) | 0.205 |
| Gemini | 35 of 36 | 985 | 65 (6.6%) | 743 (75.4%) | 0.186 |
| All four pooled | 135 | 2,623 | 666 (25.4%) | 1,496 (57.0%) | — |
The ranking is not a one-week artifact. Across the three independent weekly runs, ChatGPT's all-three share reads 8.4% → 7.9% → 5.8% and Gemini's 5.1% → 8.0% → 6.8%, while Perplexity's reads 57.8% → 76.4% → 57.2% and Claude's 30.1% → 18.0% → 23.5%. The least consistent engine never gets close to the most consistent one: in no run does ChatGPT or Gemini exceed 8.4%, and in no run does Perplexity fall below 57.2%.
What three identical asks actually look like
On 2026-08-10, ChatGPT answered "What are the best AI visibility tools for small businesses?" three times, minutes apart. The three answers cited 7, 9 and 7 domains — 19 distinct domains in total — and exactly one domain appeared in all three answers. The same day, the same engine on "What is a good Profound alternative for a small business?": 6, 7 and 6 domains per answer, 13 distinct, again exactly one present in all three. Every ChatGPT answer in the corpus carries citations (0 of 108 empty) — the sources are always there; they just change between asks. Claude is the opposite case: it declined to cite anything in 38 of its 108 answers, but when it does cite, it repeats itself more than ChatGPT does. Gemini pairs the widest net (12.3 domains per answer, 985 distinct) with the lowest repeat rate.
Why this matters for anyone measuring AI visibility
- A single-sample check on ChatGPT or Gemini is mostly noise. Roughly three quarters of the domains in any one answer will not be there if you ask twice more. "We're cited in ChatGPT" from one screenshot — or a competitor's claim that you are not — is a coin-flip observation, not a measurement.
- Perplexity screenshots are a different kind of evidence. At 63.8% all-three consistency and 0.758 mean pairwise overlap, what Perplexity cites once it mostly cites again. A presence (or absence) reading there is far more repeatable — which cuts both ways for the vendors who demo exclusively on it.
- Measure with repeated samples and score set-membership. Our own citation gauge only scores a (engine, question) pair as a citation when it holds across samples — the 2.3–2.6× within-question spread and this 9.7× engine gap are exactly why. One-shot rank trackers built for Google inherit none of these properties.
- The direction matches what week-over-week data already showed. ChatGPT's shortlist churns about twice as fast as Perplexity's on the same question across weeks, and Chris Green's 9,946-run study measured ChatGPT's domain overlap between repeated identical runs at 0.265 — within-run, our ChatGPT mean pairwise overlap reads 0.205. Same instability, three instruments.
Methodology
Within-run sample-consistency cut of the three dated four-engine production runs (2026-07-28, 2026-08-03, 2026-08-10; answer rows pulled from the production `aeoSelfAnswers` table by runId; panel version 2026-06-27, unchanged across all runs). Unit: a (question, engine, run) cell holds 3 answers from identical prompts; every URL in each answer's returned citation array is normalized to its registrable hostname (strip www — the same rule as our dated cited-domain cuts). A domain is counted "in all 3" if it appears in every sample of its cell, "in only 1" if it appears in exactly one. Cells whose three samples all returned zero citations are excluded from ratios: 0 for ChatGPT and Perplexity, 1 for Gemini, 8 for Claude (Claude returned zero citations in 38 of 108 individual answers; Gemini in 9). Percentages are per-engine sums over cells, not means of cell ratios. Full per-cell table (135 cells: sample sizes, union, all-3 and only-1 counts, pairwise Jaccard) in the cut artifact `within-run-sample-consistency-cut-2026-08-14.json`.
Does ChatGPT give the same sources every time you ask the same question?
No. Across 36 question-cells asked three times each in the same run (108 production answers over three weeks), only 7.4% of the domains ChatGPT cited appeared in all three answers to the identical question, and 75.7% appeared in just one of the three. The answer text stays on-topic; the cited sources largely reshuffle between asks.
Which AI engine is most consistent about its citations?
Perplexity, by a wide margin in our corpus: 63.8% of its cited domains appear in all three identical asks (mean pairwise overlap 0.758), versus 23.9% for Claude, 7.4% for ChatGPT and 6.6% for Gemini. The ranking held in each of the three weekly runs measured.
How many times should you ask an AI engine before trusting a visibility result?
At least three identical samples, scored on set-membership (does the domain hold across samples), not on any single answer. In our data a single ChatGPT or Gemini answer misrepresents the stable citation set roughly three times out of four; three samples separate stable citations from one-off draws.
See your number
See which businesses AI names when your client's buyers ask.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.