AI search news ·
What are IbouBot and PetalBot? A source-citing French search engine and the crawler behind Huawei's AI Search made 136 fetches of our site in 4 days
Our server-side crawler ledger — born 2026-08-17 — has already recorded two crawlers whose names appear in no AI-crawler blocklist template we have seen: IbouBot and PetalBot, 136 fetches between them by 2026-08-20. Both turn out to be search-engine fetchers with an AI-answer surface behind them. IbouBot belongs to Ibou.io, a French engine whose documentation commits, in one sentence, to the exact trade every publisher wants: "Every answer we produce cites its sources and links to them. We build the whole chain ourselves — crawling, indexing, ranking — and we do not train AI models with the data." PetalBot is Huawei's crawler; its documentation says the index it builds powers Petal Search and "content recommendations for the user in Huawei Assistant and AI Search services." Here is what each actually did on a small site, hit by hit.
Two crawlers, one week, 136 fetches
| IbouBot | PetalBot | |
|---|---|---|
| Operator (own docs) | Ibou.io — independent French search engine | Huawei (via Aspiegel) — Petal Search |
| First seen on our ledger | 2026-08-17 06:23Z | 2026-08-17 02:23Z |
| Fetches through 08-20 | 47 (46 distinct paths) | 89 (82 distinct paths) |
| Pattern | robots.txt + 1 post on day 1, then a 45-fetch full-site sweep on 08-19 in 11.8 minutes (avg 16 s between requests) | accelerating: 8 → 4 → 18 → 59 fetches/day; 59 of 89 on /news |
| Request origin (edge geo) | France, every request | Singapore, every request |
| robots.txt / sitemap.xml / llms.txt | 2 / 0 / 0 | 2 / 1 / 0 |
The instrument is our server-side crawler-hit ledger: edge middleware logs every bot-shaped user agent before any JavaScript runs, so it sees fetchers that never appear in analytics. The ledger is 3.6 days old, so "new" here means new to the instrument — PetalBot's first hit came 35 minutes after the ledger went live, so it was almost certainly crawling before we could see it. The full cut, with every count above, is saved alongside this post.
Who they are, from their own documentation
IbouBot's page describes a crawler for a search engine that "discovers publicly accessible pages and indexes them for our search engine, so that readers can find them and so that the people who published them get real visitors back." It states "We respect the robots.txt file and its disallow directives," publishes its IP ranges as JSON, supports forward-confirmed reverse DNS, promises a politeness delay of 5 seconds between queries on the same host, and is a Cloudflare Verified Bot classified as "Search." Our ledger's read matches the promises: one polite pass over the whole site — /data, /compare, /guide, /for, even the refund policy — averaging 16 seconds between requests, slower than its own stated floor.
PetalBot's page says "PetalBot is an automatic program of the Petal search engine" whose index also drives "content recommendations for the user in Huawei Assistant and AI Search services, both services are powered by Petal Search engine," and that "PetalBot complies with the Internet robots protocol." (The documentation URL inside PetalBot's own user-agent string, webmaster.petalsearch.com/site/petalbot, returned HTTP 500 when we read it; aspiegel.com — Huawei's Ireland-based operating company — hosts the working copy.) On our ledger it crawled almost entirely with its mobile user agent (87 of 89 fetches), concentrated on our news section, and fetched our sitemap once.
Why this matters if you care about AI visibility
- New citing engines index you before you have heard of them. Ibou.io's whole pitch is answers that cite and link their sources — the referral trade AI visibility depends on. Our site was swept into its index by a crawler most blocklists have never named. A blanket block-unknown-bots rule trades bandwidth you will not miss for answer surfaces that do not exist yet.
- Blocking PetalBot cuts you out of an AI surface, not just a search engine. Huawei's own documentation says the same index feeds Huawei Assistant and AI Search on a device base that classic Western SEO dashboards never report on. The popular advice to block every non-Google crawler has a price, and it is not zero.
- robots.txt is still the only control plane bots actually read. Both bots fetched robots.txt and both documentations promise to obey it. Neither touched /llms.txt — consistent with our earlier ledger read, where robots.txt outdrew llms.txt 27 fetches to 0. Across all bots in this 4-day window: 165 robots.txt fetches, 7 for llms.txt.
Whether being in Ibou's or Petal's index ever produces a citation or a visitor is a separate, measurable question — the same one we track for ChatGPT, Perplexity and Gemini. If you want to know which AI engines currently cite your business, that is what a free check reads directly.
Limits
Four days of one small site's logs. The ledger stores no IP addresses, so we did not run the reverse-DNS verification both vendors support — user-agent strings are self-declared and spoofable, though the observed politeness and origin consistency fit the declared operators. Country is edge geolocation of the connecting request. Whether either engine sends a visitor back is not yet measurable here: neither has appeared in our referral detector's lifetime log, and we will say so if that changes.
See your number
See which businesses AI names when your client's buyers ask.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.