AI search news ·

The first critical survey of GEO reviewed 45 studies and found no technique with a stable, longitudinal, cross-platform causal effect

The academic literature on generative engine optimization now has an audit. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)” by Olivier Martinez, posted to arXiv (2607.14035v1) on July 15, 2026, reviewed 45 studies published between November 16, 2023 and July 14, 2026 and graded each one by the strength of its evidence. Its central conclusion, verbatim: “already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.” Declare the interest first, as always: we sell AI-visibility audits, so this survey is a critique of the category we are in — read the paper, discount our reading of it.

What the survey did

The primary source is arXiv:2607.14035v1, dated July 15, 2026 (cs.IR, cs.DL). It searched arXiv, the ACM Digital Library, the ACL Anthology, NeurIPS, PMLR and OpenReview with backward and forward citation tracking, then sorted the resulting 45 studies into an evidence hierarchy — Level A for randomized field trials with logs and controls, down to Level E for fixed contexts and synthetic rankers. The finding that should give every buyer pause is the shape of that distribution: Level A is nearly empty. The most-cited foundational work in this field sits at the weak end, and most of what circulates as an “AEO playbook” descends from it.

The headline numbers, re-read at their source

Claim as it circulatesWhat the survey says it actually rests on
“GEO lifts visibility up to 40%”Aggarwal et al. (2024): quotation addition raised position-adjusted word count from 19.3 to 27.2 — a 41% relative gain in a fixed five-document context, not organic discoverability
“AEO multiplied our ChatGPT referrals”Watanabe & Nakayashiki (2026): treated pages rose 5.7× — but untreated pages had already risen 3.5×; the controlled time-series multiplier is 1.82 (95% CI 1.31–2.54) with a placebo test at p=0.16
“These techniques work generally”C-SEO Bench (Puerto et al. 2025): only 3 of 54 method–domain combinations were significantly positive, none in question answering, and “gains approach zero under broad adoption”
“Optimize the page body for AI”SAGEO Arena: body-only optimization reduced top-20 presence ~9%, top-10 presence 16%, and citation 6% — the opposite sign to the fixed-context result

None of this says AEO work is worthless. It says the effect sizes in circulation were measured under conditions — a supplied set of documents, a single engine, a short window — that do not describe a client's live visibility, and that the one attempt in the corpus at a controlled field measurement came back positive but tentative, with its own placebo test failing to clear significance.

The part we can corroborate from our own data

The survey's methodological core is a sentence we have been arguing from first-party measurements for six weeks: “Visibility is a distribution. It depends on the engine, date, location, query formulation, whether search is actually activated, and stochasticity in generation. A point estimate is not a stable indicator.” The evidence it cites for that: Schulte et al. (2026) measured source-overlap Jaccard of 0.34–0.42 across four engines over 45 days and recommend a minimum of 7–8 repetitions per prompt, and found 57.8% of ChatGPT repetitions did not activate web search at all; Kirsten et al. (2026) found that even at temperature zero, repeated runs change 9–28% of decisions; URL-level overlap between Google organic, AI Overviews and Gemini is 0.11–0.18, and 53% of domains cited in AI Overviews are absent from the organic top 100.

Our own run: 25 buyer questions × 4 engines × 3 repetitions, 258 completed answers, 681 brand-recommendation slots — and only 31.4% of recommended brands appeared in every repeat of the identical question on the identical engine (the full dataset). Different method, same conclusion: one answer is one draw from a distribution.

Where this cuts against us, stated plainly

Schulte's 7–8 repetitions is above what our own $79 audit does — it samples each of 25 questions three times per engine (300 answers), and the free check samples once. Three repetitions is enough to expose flicker and to report a rate instead of a verdict; it is not enough to pin a mention rate to a tight confidence interval, and we are not going to claim otherwise because an academic survey just published the number. The honest framing of any single audit — ours included — is a floor with a date and an engine attached, re-measured over time.

Why this matters for businesses and agencies

The takeaway

The first critical audit of this field's evidence base concluded that the interventions are real but narrow — they can change what happens to content the engine has already retrieved, and nobody has yet shown a technique that durably moves organic discoverability across engines. That is an argument for spending less on portable playbooks and more on repeated, per-engine measurement of the thing you are actually trying to move.

You can run the free 60-second check to see a business's mention rate across ChatGPT and Perplexity as separate numbers (the $79 audit adds Gemini and Claude, 25 questions sampled three times each — a rate, with its sample size stated). Agencies scoping AEO retainers across a book of clients can baseline five businesses at once with the $249 agency 5-pack — which is the measurement this survey says should come before any promise about what the work will move.

See your number

A free 60-second check shows what AI says about you.

Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.

Run a free AI visibility check

Begin your check

Free · 60 sec

No account · No card · 3 buyer questions, 2 engines

By running a check you agree to our Terms and Privacy Policy.

Who runs this