AI search news ·
OpenAI will now build and host your client's website. Nothing in its documentation says a search engine can read it.
On July 9, 2026 OpenAI shipped Sites in ChatGPT in public beta: describe a site in chat, ChatGPT writes it, hosts it, and gives you a URL you can share "with your team or publicly." It is a real hosting product with custom-domain support. But we read the announcement, the Sites documentation, and the help-center article on July 14, 2026, and none of them mentions indexing, crawling, robots.txt, canonical tags, or search engines even once. Publishing is documented. Being found is not. For anyone doing GEO work for clients, that gap is the entire story — and OpenAI's own flagship domain, which does publish a crawl policy, blocks Google's AI crawler, Perplexity, and Claude by name.
Primary source: OpenAI
What OpenAI shipped
The primary source is OpenAI's announcement, “ChatGPT is now a partner for your most ambitious work”, published July 9, 2026. Alongside ChatGPT Work and GPT‑5.6, OpenAI introduced Sites, in its own words: “We're also introducing Sites in ChatGPT in public beta. With Sites, you can turn your work or ideas into an interactive site or web app and share it with your team or publicly through a URL.” The Sites documentation adds that ChatGPT will “create, host, refine, and share websites, web apps, and games,” that every deployment URL is a production URL, and that a Site can be pointed at a custom domain you already own.
Two availability facts from OpenAI's help-center article matter before anything else, because they decide whether this is even usable for your book of business: “ChatGPT Sites is available in public beta on paid plans except Free and Go. It is not available in the EEA, Switzerland, or the United Kingdom at launch.” And: “Custom domains are not available in Enterprise workspaces at launch.” If your clients are UK-based, this product does not exist for them yet.
What the documentation does not say
We searched all three OpenAI pages on July 14, 2026 for the words index, crawl, robots, search engine, SEO, and discoverable. They do not appear. The docs are careful and explicit about permissions — “Hosting a Site doesn't automatically make it public,” and sharing tops out at three settings: only you or people you invite, everyone in your workspace, or anyone with the link. That is a permission model, and a good one. But “anyone with the link” is not distribution. It means nobody is stopped from visiting; it does not mean anybody can find it. A page that no crawler fetches cannot be cited by an AI engine, because there is nothing in the index to cite.
This is not a criticism of Sites for what it is built for. The use cases OpenAI lists — live dashboards, project trackers, internal portals, prototypes, interactive reports — are mostly things you would not want indexed. The failure mode is the agency that reaches for a fast site builder for a client's marketing content and assumes hosting implies reach.
Meanwhile, on the domain OpenAI does document
We fetched chatgpt.com/robots.txt on July 14, 2026. On its flagship web property, OpenAI names and fully blocks the crawlers behind the other AI engines — each with `Disallow: /`:
| Blocked crawler | Who operates it | Rule in chatgpt.com/robots.txt |
|---|---|---|
| Google-Extended | Google — the opt-out that governs Gemini grounding and training | Disallow: / |
| PerplexityBot, Perplexity‑User | Perplexity | Disallow: / |
| anthropic-ai, Claude-Web | Anthropic | Disallow: / |
| CCBot | Common Crawl — the corpus much of the AI ecosystem is built on | Disallow: / |
| Bytespider, FacebookBot, Omgili, img2dataset | ByteDance, Meta, others | Disallow: / |
For every other bot, the file is an allow-list of marketing and share paths that ends with a blanket `Disallow: /`. State the caveat plainly, because we will not overclaim: this is chatgpt.com, and OpenAI does not publish which hostname serves a deployed Site — so this is not proof that your Site is blocked. It is evidence of the platform's posture. On the one domain where OpenAI has written down a crawl policy, that policy shuts out the exact engines your client is paying you to be recommended in. Do not assume a Site inherits a friendlier one. Go and check.
What an agency should actually do
- Treat hosting and distribution as two separate purchases. A generated, hosted, shareable URL solves the build problem. It says nothing about the retrieval problem. The question to ask of any platform holding client content is not “can people open this link?” but “which crawlers are permitted to fetch it, and is it in an index?”
- Check the crawl surface before content goes on someone else's domain — it takes two minutes. Fetch `robots.txt` on whatever hostname the deployment URL actually uses. Look for a `noindex` header or meta tag on the page itself. Confirm Googlebot, Google-Extended, and PerplexityBot are not disallowed. If any of that fails, the page is a brochure you can email, not an asset that can be cited.
- Use a custom domain, which is the one path that keeps the crawl surface yours. OpenAI supports pointing a Site at an apex domain or subdomain you already own via DNS — and on a domain you control, you own the robots.txt, the canonical tag, and the indexing decision. (Not available in Enterprise workspaces at launch.) The general rule outlives this product: the durable AI-visibility asset is content on a domain you control. Everything hosted inside someone else's walled platform is theirs to make invisible.
- Re-run the client's visibility check after any migration. Moving a client's content to a new host changes the URL, the canonical, and possibly the crawl permission — three of the inputs that decide whether an engine cites them. That is a measurable before-and-after, not a vibe.
The takeaway
The AI platforms are becoming hosts. That is a genuinely useful thing — and it quietly moves the decision about who may read your client's content from your client to the platform. OpenAI has been forthright about permissions and silent about crawling, while blocking rival AI crawlers on the domain it does document. Both facts can be true at once, and an agency has to plan for both.
If you want to know whether an engine can currently name a business at all, run the free 60-second check — it reports the rate across ChatGPT and Perplexity (the $79 audit adds Gemini and Claude, 25 questions sampled three times each). If you do this for clients, the $249 agency 5-pack runs the full repeated-sample audit across five of them, white-labeled — including the before-and-after you will want the next time a client moves their content somewhere new.
See your number
A free 60-second check shows what AI says about you.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.