AI search news ·

The same buyer question, asked eight times: ChatGPT named 27 businesses, Perplexity named 11

That AI answers move is old news — we published our own version of it two weeks ago. What nobody seems to publish is which engine moves, and by how much, which is the number that decides what an agency can honestly put in a client report. Across 57 identical buyer questions that our corpus happened to ask 3–8 times each over roughly two weeks, ChatGPT's shortlist churned about twice as fast as Perplexity's on every measure we computed: 4.2% of the businesses it named survived every ask (Perplexity: 15.4%), 67.3% were named exactly once (Perplexity: 50.9%), and the mean overlap between any two asks of the same question was 0.20 against Perplexity's 0.42. Same questions, same fortnight, same extraction — one engine is roughly twice as reproducible as the other.

The accident that made this measurable

A free AskedAbout check asks an engine category questions about a market — "what are the best med spa in Charlotte, NC?" — never the business's name. Which means that when we audit several businesses in the same city and the same vertical on different days, the engines get sent the identical string several times, days or weeks apart, with no knowledge of whose audit it is. That is not an experiment we designed; it is a by-product of the roster. But it produces exactly the thing a repeat-measures test needs: 57 buyer questions asked 3–8 times each (29 on ChatGPT, 28 on Perplexity; mean 3.9 asks, mean span 15.0 and 16.6 days), inside a corpus of 173 completed audits / 1,024 answers run 2026-06-12 to 2026-07-11 (36 of the 57 questions med-spa, 11 personal-injury law, 6 dentists, 4 home services).

For each ask we reconstruct the full set of businesses the engine named — the extracted competitor set plus the audited business itself when it was named — normalise the names (lowercase, punctuation and legal suffixes stripped), and then ask a single question: of everything named in answer to this question, how much of it is still there the next time you ask? Synthetic QA rows are removed by exact-name filter before any of this, a discipline we adopted after publishing a false finding in July and correcting it in public.

Finding 1: on ChatGPT, almost nothing survives every ask

Per repeated questionChatGPTPerplexity
Businesses named per single answer6.75.2
Distinct businesses across all asks17.310.5
Named in every ask4.2% (21/502)15.4% (45/293)
Named in exactly one ask67.3% (338/502)50.9% (149/293)
Questions with no business in every ask16 of 293 of 28
Mean pairwise overlap between two asks (Jaccard)0.2040.421

Read the first two rows together, because that is where the finding lives. ChatGPT names about 6.7 businesses per answer but 17.3 different ones across four asks of the same question — it is not returning a shortlist, it is sampling from a much larger pool. Perplexity names 5.2 per answer and 10.5 across the repeats: a smaller pool, sampled more consistently. And on 16 of 29 ChatGPT questions there was not one single business that appeared in all of the asks. On Perplexity that happened 3 times in 28.

Finding 2: one question, eight asks, seventeen days

The most-repeated question in the corpus is "What are the best med spa in Charlotte, NC? Give a short ranked list with a one-line reason each." — asked 8 times on each engine between 2026-06-21 and 2026-07-08 as eight different Charlotte businesses came through the roster. Same words, same engine, 17 days:

ChatGPT — 27 distinct businesses namedAsks named in
Fanous MedSpa6 of 8
Evolve Medical Associates5 of 8
Genevieve & Co.5 of 8
Infinity MedSpa & Wellness4 of 8
Aesthetica Med Spa3 of 8
Named in all 8 asksnone
Perplexity — 11 distinct businesses namedAsks named in
Ageless Remedies8 of 8
Miramae Studio8 of 8
The Skin Center by CPS7 of 8
Voci MedSpa6 of 8
Eden MedSpa4 of 8
Named in all 8 asks2

A Charlotte med spa that is actually winning Perplexity can be identified in two asks. On ChatGPT, the best-performing business in the market missed a quarter of the asks, and the 27th name on that list appeared once and never again. Anyone who screenshots one ChatGPT answer and calls it a ranking is reporting a coin-flip that lands six ways.

Finding 3: the gap is not a name-matching artifact

The obvious objection is that ChatGPT's churn is really our extractor's churn — "Evolve Medical Associates" and "Evolve Medical Associates MedSpa" counted as two businesses would inflate the pool. So we re-ran everything with an aggressive merge that collapses any two names where one's token set contains the other's, after stripping vertical words (medspa, dental, law, group, center). Both engines improve and the gap survives intact: ChatGPT's survive-every-ask rate goes 4.2% → 8.0% and its named-once rate 67.3% → 61.4%; Perplexity goes 15.4% → 20.5% and 50.9% → 42.6%. Whatever the true rate is, ChatGPT is about 2.5× less reproducible than Perplexity on it, measured the same way on the same questions.

What this changes if you sell AI visibility to clients

Limits — what this cut does not show

Prior first-party work on the same problem, from the other direction: we re-ran 20 businesses' identical checks two weeks apart and all 20 came back different, and ChatGPT and Perplexity agree on only 10.1% of the businesses they recommend at a single point in time. This post is the third leg: the engines disagree with each other, and one of them also disagrees with itself twice as often as the other.

You can run the free 60-second check on any business to see who the engines name in its place. Agencies baselining a book of clients can run five at once with the $249 Agency 5-pack — 25 buyer questions × 4 engines × 3 samples per question, white-labeled, which exists precisely because a single ask is the thing this post says not to report.

Does this mean ChatGPT results are useless for AI visibility reporting?

No — it means they need more asks per number. On our data ChatGPT named 17.3 distinct businesses across ~4 asks of one question against Perplexity's 10.5, and only 4.2% of its names survived every ask. That is a sample-size problem, not a disqualification. Report a mention rate across repeats and state the count; do not report a rank from one answer.

How many times should I ask before a shortlist means anything?

We cannot give you a universal number and anyone who does is guessing. What our data supports: on Perplexity, the businesses that appear in every ask emerge within a handful of repeats (2 of 11 names held all 8 asks in the Charlotte question). On ChatGPT, 16 of 29 questions had no business that appeared in every ask at 3–8 repeats, so the honest answer there is that the stable set may be empty and the useful statistic is frequency, not membership.

Is this the same as your earlier post about AI answers changing?

Related but a different measurement. The earlier work re-ran the same business's own audit two weeks later and asked whether that business's result changed. This asks whether the market-level shortlist for one buyer question is stable at all, and splits the answer by engine — which is the part that turned out to be actionable, because the two engines differ by roughly 2×.

How would I reproduce this?

Pick one category-and-city question, never naming a business. Ask it on one engine on several different days, and record the full set of businesses named each time. Then compute three things: distinct names across all asks, how many appeared in every ask, and how many appeared exactly once. Run the same procedure on a second engine before concluding anything about "AI" as a category.

See your number

A free 60-second check shows what AI says about you.

Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.

Run a free AI visibility check

Begin your check

Free · 60 sec

No account · No card · 3 buyer questions, 2 engines

By running a check you agree to our Terms and Privacy Policy.

Who runs this