AI search news ·
The first critical survey of GEO reviewed 45 studies and found no technique with a stable, longitudinal, cross-platform causal effect
The academic literature on generative engine optimization now has an audit. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)” by Olivier Martinez, posted to arXiv (2607.14035v1) on July 15, 2026, reviewed 45 studies published between November 16, 2023 and July 14, 2026 and graded each one by the strength of its evidence. Its central conclusion, verbatim: “already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.” Declare the interest first, as always: we sell AI-visibility audits, so this survey is a critique of the category we are in — read the paper, discount our reading of it.
Primary source: arXiv — Martinez, “A Critical Survey of Generative Engine Optimization”
What the survey did
The primary source is arXiv:2607.14035v1, dated July 15, 2026 (cs.IR, cs.DL). It searched arXiv, the ACM Digital Library, the ACL Anthology, NeurIPS, PMLR and OpenReview with backward and forward citation tracking, then sorted the resulting 45 studies into an evidence hierarchy — Level A for randomized field trials with logs and controls, down to Level E for fixed contexts and synthetic rankers. The finding that should give every buyer pause is the shape of that distribution: Level A is nearly empty. The most-cited foundational work in this field sits at the weak end, and most of what circulates as an “AEO playbook” descends from it.
The headline numbers, re-read at their source
| Claim as it circulates | What the survey says it actually rests on |
|---|---|
| “GEO lifts visibility up to 40%” | Aggarwal et al. (2024): quotation addition raised position-adjusted word count from 19.3 to 27.2 — a 41% relative gain in a fixed five-document context, not organic discoverability |
| “AEO multiplied our ChatGPT referrals” | Watanabe & Nakayashiki (2026): treated pages rose 5.7× — but untreated pages had already risen 3.5×; the controlled time-series multiplier is 1.82 (95% CI 1.31–2.54) with a placebo test at p=0.16 |
| “These techniques work generally” | C-SEO Bench (Puerto et al. 2025): only 3 of 54 method–domain combinations were significantly positive, none in question answering, and “gains approach zero under broad adoption” |
| “Optimize the page body for AI” | SAGEO Arena: body-only optimization reduced top-20 presence ~9%, top-10 presence 16%, and citation 6% — the opposite sign to the fixed-context result |
None of this says AEO work is worthless. It says the effect sizes in circulation were measured under conditions — a supplied set of documents, a single engine, a short window — that do not describe a client's live visibility, and that the one attempt in the corpus at a controlled field measurement came back positive but tentative, with its own placebo test failing to clear significance.
The part we can corroborate from our own data
The survey's methodological core is a sentence we have been arguing from first-party measurements for six weeks: “Visibility is a distribution. It depends on the engine, date, location, query formulation, whether search is actually activated, and stochasticity in generation. A point estimate is not a stable indicator.” The evidence it cites for that: Schulte et al. (2026) measured source-overlap Jaccard of 0.34–0.42 across four engines over 45 days and recommend a minimum of 7–8 repetitions per prompt, and found 57.8% of ChatGPT repetitions did not activate web search at all; Kirsten et al. (2026) found that even at temperature zero, repeated runs change 9–28% of decisions; URL-level overlap between Google organic, AI Overviews and Gemini is 0.11–0.18, and 53% of domains cited in AI Overviews are absent from the organic top 100.
Our own run: 25 buyer questions × 4 engines × 3 repetitions, 258 completed answers, 681 brand-recommendation slots — and only 31.4% of recommended brands appeared in every repeat of the identical question on the identical engine (the full dataset). Different method, same conclusion: one answer is one draw from a distribution.
Where this cuts against us, stated plainly
Schulte's 7–8 repetitions is above what our own $79 audit does — it samples each of 25 questions three times per engine (300 answers), and the free check samples once. Three repetitions is enough to expose flicker and to report a rate instead of a verdict; it is not enough to pin a mention rate to a tight confidence interval, and we are not going to claim otherwise because an academic survey just published the number. The honest framing of any single audit — ours included — is a floor with a date and an engine attached, re-measured over time.
Why this matters for businesses and agencies
- Treat any AEO uplift promise as un-evidenced until it names its level. The survey's grading is the useful buyer's tool: ask whether a claimed result came from a randomized field trial with logs (Level A) or from a fixed-document experiment (Level E). Most vendor decks — and most of the “40%” slides — are quoting Level E as though it were Level A.
- “Transfers poorly” is the operative phrase. Only topical relevance and positioning replicated across the corpus; general heuristics did not. A checklist that worked for one client's category is a hypothesis for the next one, not a deliverable.
- Visibility is indexed by engine and surface. With 0.11–0.18 URL overlap between Google organic, AI Overviews and Gemini, a rank report and an AI-citation report are measurements of different things. Reporting one as if it covers the other is the most common measurement error we see.
- Under broad adoption, gains compress. C-SEO Bench's “gains approach zero under broad adoption” is the strategic caveat: the techniques being sold today assume most competitors are not doing them. Measurement of where you actually stand keeps its value either way; the tactic-of-the-month does not.
The takeaway
The first critical audit of this field's evidence base concluded that the interventions are real but narrow — they can change what happens to content the engine has already retrieved, and nobody has yet shown a technique that durably moves organic discoverability across engines. That is an argument for spending less on portable playbooks and more on repeated, per-engine measurement of the thing you are actually trying to move.
You can run the free 60-second check to see a business's mention rate across ChatGPT and Perplexity as separate numbers (the $79 audit adds Gemini and Claude, 25 questions sampled three times each — a rate, with its sample size stated). Agencies scoping AEO retainers across a book of clients can baseline five businesses at once with the $249 agency 5-pack — which is the measurement this survey says should come before any promise about what the work will move.
See your number
A free 60-second check shows what AI says about you.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.