AI search news ·

Where do AI crawlers actually go on a small site? In 7 days ChatGPT-User fetched our homepage 193 times (71% of its requests), Googlebot spent 56% of its requests on robots.txt and sitemap.xml, and Amazonbot read 151 distinct pages — more than any AI fetcher

Between 2026-08-19 13:53 UTC and 2026-08-26 13:37 UTC our crawler ledger logged 1,955 bot-shaped requests to askedabout.com; ChatGPT-User made 272 of them and 193 (71%) were for the homepage. Googlebot made 176 requests and 99 (56%) were robots.txt or sitemap.xml, leaving 71 fetches of 45 distinct content pages on a site with 151 live URLs. Amazonbot, which no AI answer engine of ours has ever cited through, fetched 151 distinct content pages — more than PerplexityBot (67), ClaudeBot (66) or OAI-SearchBot (57). Perplexity-User and Google-Extended made 0 requests. Method, the full per-fetcher allocation table, the 728 requests from bots our classifier does not name, and the pre-committed read for next Wednesday are below.

Method, stated before the numbers

Every request to askedabout.com whose User-Agent looks like a bot (matches `bot|crawl|spider|slurp|fetch|-user|…`) is forwarded by the site's edge middleware to a Convex table, with the path, the UA, and the edge's country code. A fixed token table then names the fetcher by substring — `ChatGPT-User`, `OAI-SearchBot`, `GPTBot`, `PerplexityBot`, `Perplexity-User`, `ClaudeBot`, `Claude-User`, `Amazonbot`, `Googlebot`, `Google-Extended`, `bingbot` and about a dozen more. A bot-shaped UA that matches no token is stored as "other" with its UA intact, never dropped. We described the ledger and its fetch-speed read in an earlier post.

This cut pulls the trailing seven days (2026-08-19 13:53 to 2026-08-26 13:37 UTC, 6.99 days) — 2,035 rows. Two kinds of row are excluded and counted: 58 are our own healthcheck's labelled probe requests, and 22 are requests for secret-shaped paths (`/.env`, `/.git/HEAD`, `/etc/passwd`…) made under AI-crawler User-Agents, which we showed on 2026-08-22 are a vulnerability scanner wearing the badge. That leaves 1,955 net requests across 172 distinct paths: 1,227 from fetchers the token table names and 728 from bots it does not.

Each path is then assigned one class by a fixed rule, in this order: the homepage (`/`); plumbing (`/robots.txt`, `/sitemap.xml`, `/llms.txt`, `/favicon.ico`, `.well-known`, `humans.txt`, `ai.txt`); a `/news/` post; the `/news` index; the `/compare`, `/guide`, `/data`, `/for` and `/library` collections; a site page (`/about`, `/terms`, `/privacy-policy`, `/refund-policy`, `/pricing`, `/agency-5-pack`…); anything else is unclassified (11 requests, all listed in the artifact). "Content pages" below means everything except the homepage, plumbing and unclassified. A hit is one HTTP request; nothing is de-duplicated by IP, and — this matters — a User-Agent that claims to be Googlebot is counted as Googlebot unless it asked for a secret-shaped path. The classifier, the class rule and every count are in the linked script and artifact.

The allocation, fetcher by fetcher

FetcherRequestsDistinct pathsHomepagePlumbing (robots/sitemap)/news posts (distinct)/compare + /guide + /data + /for + /libraryAll content pages
ChatGPT-User (OpenAI, user-directed)27235193 (71%)0 (0%)59 (23 posts)1879 (29%; 34 distinct)
Googlebot176486 (3%)99 (56%)39 (24 posts)2671 (40%; 45 distinct)
Amazonbot1641552 (1%)0 (0%)108 (105 posts)42159 (97%; 151 distinct)
ClaudeBot (Anthropic, index)139691 (1%)64 (46%)44 (44 posts)2374 (53%; 66 distinct)
OAI-SearchBot (OpenAI, search index)129604 (3%)46 (36%)59 (41 posts)1778 (60%; 57 distinct)
PerplexityBot126694 (3%)11 (9%)43 (31 posts)59111 (88%; 67 distinct)
bingbot124524 (3%)29 (23%)73 (35 posts)1491 (73%; 48 distinct)
YandexBot552110 (18%)25 (45%)20 (18 posts)020 (36%; 18 distinct)
Claude-User (Anthropic, user-directed)2560 (0%)11 (44%)14 (5 posts)014 (56%; 5 distinct)
GPTBot (OpenAI, training)821 (13%)7 (88%)0 (0 posts)00 (0%; 0 distinct)
DuckDuckBot · Applebot · DuckAssistBot · meta-externalagent5 · 2 · 1 · 12 · 1 · 1 · 14 · 0 · 0 · 10 · 2 · 1 · 01 · 0 · 0 · 001 · 0 · 0 · 0
Perplexity-User · Google-Extended · Bytespider · CCBot0
Other bot-shaped (unnamed by the classifier)72816075 (10%)135 (19%)311 (106 posts)170511 (70%; 151 distinct)

Three things in that table are the finding. First, ChatGPT-User — the fetcher OpenAI uses when a live ChatGPT answer opens a page on a user's behalf — is almost entirely a homepage reader here: 193 of 272 requests (71%), 0 plumbing, and 59 requests spread over 23 of our 108 fetched posts; the most-fetched post drew 10. On this site the homepage is also where 9 of 9 ChatGPT citations of us have landed across four weekly gauge runs, so the fetch pattern and the citation pattern agree. Second, Googlebot's week was mostly bookkeeping: 99 of 176 requests were robots.txt (80) or sitemap.xml (19), and the remaining 71 content fetches touched 45 of 151 live URLs — 24 of them /news posts, in a week when we published 8 and had 20 URLs sitting unindexed. Third, the widest reader is not an answer engine: Amazonbot fetched 151 distinct content pages including 105 of the 108 posts, one request each, 97% content and 0% plumbing, against PerplexityBot's 67, ClaudeBot's 66 and OAI-SearchBot's 57.

Two smaller readings. ClaudeBot's 139 requests came on 4 of the 8 calendar days, 97 of them on Monday 2026-08-24, and exactly half were the sitemap (32) and robots.txt (32) — it re-reads the map, then reads 44 posts once each. GPTBot, OpenAI's training crawler, made 8 requests: 7 for robots.txt and 1 for the homepage, 0 content pages — on this site the OpenAI reader that matters is OAI-SearchBot (129) and the OpenAI reader that matters most is ChatGPT-User.

The 728 requests from bots the classifier does not name

UA familyRequestsDistinct paths
PetalBot203148
AhrefsBot10474
SemrushBot9659
IbouBot5754
SEBot-WA4417
HubSpot377
SleepBot2927
MJ12bot216
DotBot2018
ShapBot168
GEO-Brain-CitationBot127
PipericBot126
26 further families77

37% of the week's bot requests came from fetchers our table does not name, and they are a different population: SEO-tool crawlers (AhrefsBot, SemrushBot, MJ12bot, DotBot), Huawei's PetalBot (203 requests across 148 paths, the single busiest reader of the week on any list), IbouBot, and a tail of small "AI citation" crawlers (GEO-Brain-CitationBot, CitableBot, SpyglassesBot, ShapBot) that fetch a handful of pages each. These read content pages 70% of the time. None of them is an engine that can cite a page to a buyer; all of them are in the denominator of any "AI crawler traffic" number that counts bot-shaped User-Agents without naming them.

Where the week's requests went, all fetchers

Path classAll bot requestsFrom named fetchersShare of named
homepage30523019%
plumbing (robots, sitemap, llms.txt…)43029524%
/news posts (108 distinct paths)77146037%
/compare133776%
/guide98494%
/data76423%
/for55282%
site pages + /library + /news index76423%
unclassified1140%
total1,9551,227100%

Named fetchers put 24% of their requests into plumbing and 19% into the homepage before any content page is read; the /news class, which holds 107 of the site's 151 live URLs, drew 37%. The four commercial classes together (`/compare`, `/guide`, `/data`, `/for`, 37 live URLs) drew 16% — 196 named-fetcher requests, of which PerplexityBot alone made 59.

Limits

What we will read next Wednesday

The same cut over the next seven days (2026-08-26 13:37 to 09-02, run on our 09-02 growth run) is pre-committed on three numbers. (A) ChatGPT-User's homepage share stays at or above 50% of its requests, else "ChatGPT reads the homepage" is falsified for this site; below 30% we say so at the top. (B) Googlebot's plumbing share stays at or above 40%, else the "mostly bookkeeping" reading is falsified. (C) Amazonbot's distinct content paths remain at or above each of PerplexityBot's, ClaudeBot's and OAI-SearchBot's, else "the widest reader is not an answer engine" is falsified as stated. Each is a number in next week's artifact, not an interpretation.

The script is `content/news/crawler_allocation_cut.mjs`; the artifact is `content/news/crawler-allocation-cut-2026-08-26.json` (9 of 9 self-checks: the probe/spoofed/net partition, every fetcher's classes summing to its requests, class and per-day totals summing to the net, UA families summing to the unnamed bucket), and the slimmed row dump (timestamp, vendor, token, kind, path, country, UA family — no full UA strings) is `content/news/crawler-rows-slim-2026-08-19-to-26.json`. If you run a small site and want the same read, the ledger is 120 lines of middleware and one table. If you would rather know what the engines say about your business than which of their crawlers visited, the check on our homepage asks them.

Does ChatGPT-User fetching a page mean ChatGPT cited it?

No. ChatGPT-User is the fetcher a live answer uses to open a page; whether the page is then cited, mentioned or discarded is not visible from the server. On our site the fetch pattern (71% homepage) and the citation pattern (9 of 9 ChatGPT citations on the homepage across four weekly gauge runs) agree, which is consistent with fetch-then-cite, not proof of it.

Why does Googlebot spend so much of its budget on robots.txt and sitemap.xml?

Googlebot re-reads robots.txt before crawl sessions and sitemap.xml when it is resubmitted or has a fresh lastmod; a small site that resubmits its sitemap daily and publishes daily will see many such reads. The number that matters for indexing is the other 40%: 71 content fetches on 45 distinct URLs in a week when 20 URLs were unindexed.

Should I block Amazonbot or the SEO-tool crawlers to save AI crawl budget?

Not on this evidence. There is no shared budget between fetchers: blocking Amazonbot does not make OAI-SearchBot fetch more. Block a crawler only if you do not want its operator's products to read your pages; Amazonbot feeds Alexa and Amazon's search features, and the SEO-tool crawlers feed the backlink indexes your competitors' tools read.

See your number

See which businesses AI names when your client's buyers ask.

Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.

Check a client's AI visibility

Begin your check

Free · 60 sec

No account · No card · 3 buyer questions, 2 engines

By running a check you agree to our Terms and Privacy Policy.

Who runs this