AI search news ·
The same buyer question, asked eight times: ChatGPT named 27 businesses, Perplexity named 11
That AI answers move is old news — we published our own version of it two weeks ago. What nobody seems to publish is which engine moves, and by how much, which is the number that decides what an agency can honestly put in a client report. Across 57 identical buyer questions that our corpus happened to ask 3–8 times each over roughly two weeks, ChatGPT's shortlist churned about twice as fast as Perplexity's on every measure we computed: 4.2% of the businesses it named survived every ask (Perplexity: 15.4%), 67.3% were named exactly once (Perplexity: 50.9%), and the mean overlap between any two asks of the same question was 0.20 against Perplexity's 0.42. Same questions, same fortnight, same extraction — one engine is roughly twice as reproducible as the other.
Primary source: AskedAbout (first-party audit corpus, 173 audits / 1,024 answers)
The accident that made this measurable
A free AskedAbout check asks an engine category questions about a market — "what are the best med spa in Charlotte, NC?" — never the business's name. Which means that when we audit several businesses in the same city and the same vertical on different days, the engines get sent the identical string several times, days or weeks apart, with no knowledge of whose audit it is. That is not an experiment we designed; it is a by-product of the roster. But it produces exactly the thing a repeat-measures test needs: 57 buyer questions asked 3–8 times each (29 on ChatGPT, 28 on Perplexity; mean 3.9 asks, mean span 15.0 and 16.6 days), inside a corpus of 173 completed audits / 1,024 answers run 2026-06-12 to 2026-07-11 (36 of the 57 questions med-spa, 11 personal-injury law, 6 dentists, 4 home services).
For each ask we reconstruct the full set of businesses the engine named — the extracted competitor set plus the audited business itself when it was named — normalise the names (lowercase, punctuation and legal suffixes stripped), and then ask a single question: of everything named in answer to this question, how much of it is still there the next time you ask? Synthetic QA rows are removed by exact-name filter before any of this, a discipline we adopted after publishing a false finding in July and correcting it in public.
Finding 1: on ChatGPT, almost nothing survives every ask
| Per repeated question | ChatGPT | Perplexity |
|---|---|---|
| Businesses named per single answer | 6.7 | 5.2 |
| Distinct businesses across all asks | 17.3 | 10.5 |
| Named in every ask | 4.2% (21/502) | 15.4% (45/293) |
| Named in exactly one ask | 67.3% (338/502) | 50.9% (149/293) |
| Questions with no business in every ask | 16 of 29 | 3 of 28 |
| Mean pairwise overlap between two asks (Jaccard) | 0.204 | 0.421 |
Read the first two rows together, because that is where the finding lives. ChatGPT names about 6.7 businesses per answer but 17.3 different ones across four asks of the same question — it is not returning a shortlist, it is sampling from a much larger pool. Perplexity names 5.2 per answer and 10.5 across the repeats: a smaller pool, sampled more consistently. And on 16 of 29 ChatGPT questions there was not one single business that appeared in all of the asks. On Perplexity that happened 3 times in 28.
Finding 2: one question, eight asks, seventeen days
The most-repeated question in the corpus is "What are the best med spa in Charlotte, NC? Give a short ranked list with a one-line reason each." — asked 8 times on each engine between 2026-06-21 and 2026-07-08 as eight different Charlotte businesses came through the roster. Same words, same engine, 17 days:
| ChatGPT — 27 distinct businesses named | Asks named in |
|---|---|
| Fanous MedSpa | 6 of 8 |
| Evolve Medical Associates | 5 of 8 |
| Genevieve & Co. | 5 of 8 |
| Infinity MedSpa & Wellness | 4 of 8 |
| Aesthetica Med Spa | 3 of 8 |
| Named in all 8 asks | none |
| Perplexity — 11 distinct businesses named | Asks named in |
|---|---|
| Ageless Remedies | 8 of 8 |
| Miramae Studio | 8 of 8 |
| The Skin Center by CPS | 7 of 8 |
| Voci MedSpa | 6 of 8 |
| Eden MedSpa | 4 of 8 |
| Named in all 8 asks | 2 |
A Charlotte med spa that is actually winning Perplexity can be identified in two asks. On ChatGPT, the best-performing business in the market missed a quarter of the asks, and the 27th name on that list appeared once and never again. Anyone who screenshots one ChatGPT answer and calls it a ranking is reporting a coin-flip that lands six ways.
Finding 3: the gap is not a name-matching artifact
The obvious objection is that ChatGPT's churn is really our extractor's churn — "Evolve Medical Associates" and "Evolve Medical Associates MedSpa" counted as two businesses would inflate the pool. So we re-ran everything with an aggressive merge that collapses any two names where one's token set contains the other's, after stripping vertical words (medspa, dental, law, group, center). Both engines improve and the gap survives intact: ChatGPT's survive-every-ask rate goes 4.2% → 8.0% and its named-once rate 67.3% → 61.4%; Perplexity goes 15.4% → 20.5% and 50.9% → 42.6%. Whatever the true rate is, ChatGPT is about 2.5× less reproducible than Perplexity on it, measured the same way on the same questions.
What this changes if you sell AI visibility to clients
- Stop reporting a position, start reporting a rate. "You are #3 on ChatGPT" describes one draw from a pool of 17. "You were named in 6 of 8 asks" is a measurement. The second one survives the client asking you to prove it in the meeting.
- Sample count is engine-specific. On our data a Perplexity shortlist stabilises in a handful of asks; a ChatGPT one does not stabilise at anything we can afford to run. Budget repeats where the variance actually is, and say in the report how many asks the number came from.
- A client's "but I checked and I wasn't there" is usually not a regression. With 67.3% of ChatGPT names appearing exactly once, absence from a single answer is the base case, not a data point. Having the distribution in front of you is what turns that call from a defensive one into a routine one.
- Engine choice is a reporting decision, not just an audience decision. If you need a number a client can track month over month, the engine that repeats itself is worth more per dollar of measurement than the one with more users — as long as you disclose which one you measured.
Limits — what this cut does not show
- These are natural repeats, not a controlled experiment. The asks are separated by days or weeks, so this mixes the model's own sampling variance with real change in what it retrieves. We cannot separate those two — but a client cannot either, which is the practical point. A within-minute repeat would isolate sampling noise and we have not run one at this scale.
- Free-check engines only. The free check runs ChatGPT + Perplexity (`gpt-5-search-api` and `sonar`). Gemini and Claude are in the paid audit and are not in this cut, so nothing here should be read as a claim about them.
- Names come from a per-answer extraction step, so both engines carry extraction noise. Finding 3 bounds it; it does not eliminate it.
- Med-spa weighted. 36 of the 57 repeated questions are med spas, because that is the densest part of our roster. The per-engine direction is the same in the other three verticals, but they are thin.
- Convenience sample. These are businesses we chose to audit, in US metros. Directionally strong, not nationally representative.
- We sell this measurement. AskedAbout sells AI-visibility audits, so discount the conclusion and check the method — which is why the method is stated in full and the counts are printed.
Prior first-party work on the same problem, from the other direction: we re-ran 20 businesses' identical checks two weeks apart and all 20 came back different, and ChatGPT and Perplexity agree on only 10.1% of the businesses they recommend at a single point in time. This post is the third leg: the engines disagree with each other, and one of them also disagrees with itself twice as often as the other.
You can run the free 60-second check on any business to see who the engines name in its place. Agencies baselining a book of clients can run five at once with the $249 Agency 5-pack — 25 buyer questions × 4 engines × 3 samples per question, white-labeled, which exists precisely because a single ask is the thing this post says not to report.
Does this mean ChatGPT results are useless for AI visibility reporting?
No — it means they need more asks per number. On our data ChatGPT named 17.3 distinct businesses across ~4 asks of one question against Perplexity's 10.5, and only 4.2% of its names survived every ask. That is a sample-size problem, not a disqualification. Report a mention rate across repeats and state the count; do not report a rank from one answer.
How many times should I ask before a shortlist means anything?
We cannot give you a universal number and anyone who does is guessing. What our data supports: on Perplexity, the businesses that appear in every ask emerge within a handful of repeats (2 of 11 names held all 8 asks in the Charlotte question). On ChatGPT, 16 of 29 questions had no business that appeared in every ask at 3–8 repeats, so the honest answer there is that the stable set may be empty and the useful statistic is frequency, not membership.
Is this the same as your earlier post about AI answers changing?
Related but a different measurement. The earlier work re-ran the same business's own audit two weeks later and asked whether that business's result changed. This asks whether the market-level shortlist for one buyer question is stable at all, and splits the answer by engine — which is the part that turned out to be actionable, because the two engines differ by roughly 2×.
How would I reproduce this?
Pick one category-and-city question, never naming a business. Ask it on one engine on several different days, and record the full set of businesses named each time. Then compute three things: distinct names across all asks, how many appeared in every ask, and how many appeared exactly once. Run the same procedure on a second engine before concluding anything about "AI" as a category.
See your number
A free 60-second check shows what AI says about you.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.