AI search news ·

Does ChatGPT read your page before it cites it? In 1,249 captured answers, 1 page in 80 was opened — and an opened page is cited 74% of the time against 7% for the rest

When ChatGPT answers with the web, it almost never reads your page. RESONEO's reverse-engineering study — 1,249 real ChatGPT answers captured in July 2026 and re-run across every tier and mode in August — finds that the model works from a title, a URL and roughly two hundred characters of snippet fetched by one of four retrieval engines, and opens the page itself in about one case in 80, historically only in paid thinking mode. Of 61,332 URLs that appeared in answers' Sources panels, 759 were opened and read (1.2%); 757 of those opens sat in paid thinking mode. When a page is opened, it is cited 74% of the time; a page merely present in the grounding set is cited 7% of the time. Which engine fetched the snippet depends on the mode more than the subscription: free instant mode drew 63.6% of search results from OpenAI's own index (internally named `labrador`), paid thinking mode drew 74.6% from scraped Google, and the August free "Think" button runs 74.7% on labrador while paid thinking stays 75.3% on Google. The labrador snippet is a fixed cut of about 202 characters taken from the start of your rendered body, anchored on the H1, with the meta description ignored — so the study's advice reduces to one line: your grounding budget is your full title plus about 150 useful characters after your H1. We checked the study's homepage hypothesis against our own 144-answer run from this morning: ChatGPT cited a homepage in 26.5% of its 279 citations, against 1.1% for Perplexity, 1.9% for Claude and 0% for Gemini — and all four citations of AskedAbout that ChatGPT produced this week were `askedabout.com/`, not a deep page.

What RESONEO measured, and how

The study is published on RESONEO's research site as "What ChatGPT pulls, what it shows, what it cites" (July 2026, updated August 2026; its co-founder Olivier de Segonzac summarized it for Search Engine Land on 2026-08-17, with the conflict disclosed). The method is network capture, not answer reading: a Chrome extension recorded the raw data stream behind 1,249 real ChatGPT answers (682 instant, 567 thinking; just under 500 on free accounts, just over 750 on paid) across a fleet of accounts and countries, yielding roughly 88,000 search results and 26,900 distinct pages from 6,400 domains. Until 2026-07-21 the stream carried an undocumented `result_source` field naming which engine returned each result (760 of the 1,249 answers carry it; the field went bare on 07-22 and the team re-identified engines from formatting signatures afterwards). RESONEO also ran canary pages on its own domain with server logs, which is how the cache-retention and crawl-discovery findings below are grounded in log lines rather than inference. It credits earlier network-traffic work by Suganthan Mohanadasan and others as the starting point.

The four engines behind a ChatGPT answer

Engine (internal name)What it isFree instant mode (search results only, 1,793 results / 158 answers)Paid thinking mode (16,407 results / 176 answers)
labradorOpenAI's own in-house index; snippet cut at 202 chars, anchored on the H163.6%23.6%
brightScraped Google (RESONEO found the string `shopping_brightdata_new_feed_serving_week_2026_06_22` in OpenAI's own code); 1 URL in 3 on Google page 1, 9 of 10 titles match Google character-for-character34.7%74.6%
oxylabsScraped Google, news results only0% of search results (12.7% of instant mode when news is counted)
serpFourth pipeline, mostly news1.6%1.9%

Two findings sit behind that table. First, labrador is neither rebadged Bing nor a licensed-publisher catalogue: Bing caps displayed titles at 75 characters and a quarter of labrador titles run past that (the longest seen: 977), only one labrador URL in twenty appears on Bing's first page, and outlets with no OpenAI licensing deal receive exactly the same format, length and freshness as partners. RESONEO never observed a `bing` engine on any of its accounts and hypothesises it is exposed only to Team plans. Second, the split depends on the question, not just the tier: on stable questions labrador's free-tier share rises to 100%; on business-and-place queries it was 100% across 310 snippets, product 97.6%, news 52.6% labrador / 46.1% Google. The August update re-ran the July prompts across every configuration: the free "Think" button pulls as many URLs per conversation as paid thinking (35.3 vs 35.1) but from the opposite corpus — 74.7% labrador on free Think, 75.3% scraped Google on paid thinking. "A good Google position still helps, but it does not guarantee visibility on the free Think path."

The pyramid: pulled, promoted, cited, opened

Layer (across the 1,249 answers)CountNote
Every URL in the Sources side panel61,33253,716 (≈9 in 10) stay at the bottom under "More", never promoted or cited
Attached to a citation tooltip7,616roughly a third are attached without being the named lead source ("+1")
First citation visible in the tooltip5,032the pill the reader actually sees
Opened and read by the model7591.2% of side-panel URLs; 757 in paid thinking mode; only 2 free-tier conversations opened a page (July)

The median answer lists about twenty pages at the bottom of the panel, promotes five to the top, cites three in its text and opens none — nine answers in ten open nothing. But opening is decisive: on the 440 opened pages exposed as result objects, 326 were cited (74%), against 4,303 of 57,853 merely-retrieved pages (7%). What gets opened is specific: regulatory and official material, primary sources, pricing pages, newsrooms — "pages deep in the tree, almost never home pages" (most-opened domains: openai.com, anthropic.com, learn.microsoft.com, digital-strategy.ec.europa.eu). The August update lifts the paid lock slightly: with the free Think button, a free account now opens 0.28 pages per conversation. Two mechanics matter for anyone reading their analytics. Instant mode cites from memory a lot: 85% of instant-mode citations point to a URL that appears nowhere in the retrieval net (31% in thinking), some from widgets, "the rest looks like brand homepages written from the model's own memory." And ChatGPT appends `utm_source=chatgpt.com` to 95% of displayed links (every one in instant mode) — but never to the pages the model opened and read, so "the citations worth the most … are precisely those your analytics tool will fail to attribute to ChatGPT."

What the model actually sees of your page: 202 characters, from the H1

This is the operational half. The labrador snippet is not query-dependent and not built at search time — seven regional variants of the same press release, queried identically, produce seven different anchors; pages seen under several queries return an identical snippet. It is frozen at crawl time, cut just past 200 characters, and taken from the start of the rendered body, not from metadata: your meta description is ignored (the Google scrape picks it up one time in three). RESONEO fetched 534 pages ChatGPT had actually cited and compared them with the stored snippet: of the 463 with an H1, 387 snippets contain it (83.6%), and in 297 of those the H1 is the very first character (77%); when there is a prefix it is 20 characters at the median. One page in seven has no H1 markup at all, at which point the anchor becomes unpredictable. Hence the study's one-line advice: "Your grounding budget in instant mode is your full title plus roughly one hundred and fifty useful characters starting at your H1. Mark up an H1, clear the runway between it and the first paragraph." Three more limits are worth writing down: pages over four megabytes are not read at all ("no partial read, no truncation"); pages reach the model converted to Markdown, with JSON-LD stripped on an on-demand read (RESONEO is careful to say this says nothing about what the index does with structured data); and there is a two-layer cache — a page is not re-fetched within 30 minutes, and cached copies were still being served 91 days after the last logged OpenAI robot visit, with a server log to back it. On discovery: GPTBot downloaded RESONEO's sitemap roughly once a day and followed none of the eight URLs listed only there, while crawling all eight pages reachable through a link the same morning; Bingbot, ClaudeBot and GoogleOther did the exact opposite. "At OpenAI, discovery runs on links."

What our own run this morning adds: the homepage share, engine by engine

RESONEO's "brand homepages written from memory" reading is a hypothesis it flags as untestable from its data. Our weekly self-audit is a different instrument on a different path — 12 small-business buyer questions × 4 engines × 3 samples through the vendors' APIs (OpenAI's web-search tool, not the ChatGPT client; every one of the 279 OpenAI citation URLs this week carries `?utm_source=openai`, not `chatgpt.com`) — but it stores every cited URL, so we cut this morning's run (2026-08-17, 13:11 UTC, run `jh72qsgx…`, 144 of 144 answers completed) by URL depth. ChatGPT cited a homepage (path `/`) in 74 of its 279 citations, 26.5%, in 20 of 36 answers, across 54 distinct domains. Perplexity: 8 of 710 (1.1%). Claude: 5 of 266 (1.9%). Gemini: 0 of 382. That is the fourth consecutive weekly run reproducing the outlier we measured on the previous three (28.4% across 07-28/08-03/08-10), and it is consistent with — not proof of — RESONEO's reading that ChatGPT's instant-style citations lean on remembered brand roots while the other engines link the page that answered. Our own four ChatGPT citations this week are the same shape: `askedabout.com/` in all three samples of "How can I check what ChatGPT says about my business?" and one sample of "Which tools show how ChatGPT, Perplexity and Gemini recommend businesses?" — the homepage, never a /news or /data page, while Perplexity's three citations of us were all a deep comparison page. Applying the study's snippet rule to the two pages of ours that were cited this week: our homepage's H1 is 39 characters and the 202-character window from it reads "What is AI saying about your business? Your customers now ask ChatGPT, Perplexity and Gemini who to buy from. A free sixty-second check shows whether AI recommends you — or quietly sends them elsewhere"; the comparison page's H1 is 69 characters and its window is the price sentence. Both pages are 64 KB and 97 KB — a long way from the 4 MB wall. One more line from our own logs, in the RESONEO spirit of saying what we can see: in the ~14 hours of today's crawler window, OAI-SearchBot fetched 51 distinct paths on askedabout.com and `/robots.txt` three times, GPTBot fetched exactly one page (the homepage), and no OpenAI agent requested `/sitemap.xml` once. Artifact: `content/news/chatgpt-homepage-share-run3-2026-08-17.json`.

Why this matters if you care about AI visibility

Does ChatGPT read a web page before citing it?

Rarely. In RESONEO's capture of 1,249 real ChatGPT answers (July 2026), 759 pages were opened and read across 61,332 side-panel URLs — 1.2%, about one page in 80 — and 757 of those opens were in paid thinking mode. Nine answers in ten open nothing. When a page is opened it is cited 74% of the time; a page that is merely retrieved is cited 7% of the time.

Does ChatGPT search use Google or Bing?

Both a scraped copy of Google and OpenAI's own index, per RESONEO's network capture: free instant mode drew 63.6% of search results from OpenAI's in-house index (internally 'labrador') and 34.7% from scraped Google ('bright'); paid thinking mode drew 74.6% from scraped Google. RESONEO never observed a Bing pipeline on its accounts and rules out labrador being rebadged Bing (a quarter of its titles exceed Bing's 75-character cap).

Does ChatGPT use my meta description?

Not on OpenAI's own index. RESONEO found the labrador snippet is about 202 characters taken from the start of the rendered body, anchored on the H1 (present in 83.6% of 463 checked snippets, first character in 77%), and that the meta description is ignored; the Google-scrape pipeline picks the meta description up about one time in three.

Do AI engines cite homepages or deep pages?

Mostly deep pages, except ChatGPT. In our 2026-08-17 run of 12 buyer questions × 4 engines × 3 samples, ChatGPT cited a homepage in 26.5% of its 279 citations; Perplexity 1.1%, Claude 1.9%, Gemini 0%. Across our previous three weekly runs the ChatGPT figure was 28.4% and 92.3% of all cited URLs were deep pages.

See your number

See which businesses AI names when your client's buyers ask.

Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.

Check a client's AI visibility

Begin your check

Free · 60 sec

No account · No card · 3 buyer questions, 2 engines

By running a check you agree to our Terms and Privacy Policy.

Who runs this