AI search news ·
Google is swapping the model underneath Search — and no AI-visibility baseline survives a generator change
On 21 July 2026 Google announced three models — Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber — and put one sentence in the launch post that matters more to AI visibility than the benchmark tables: "3.5 Flash-Lite is also rolling out in Google Search." Flash-Lite is the cheap, fast tier: $0.30 per million input tokens, $2.50 per million output, 350 output tokens/second, explicitly built for "agentic search and document processing." The measurement consequence is unglamorous and immediate: any number you recorded about how Google's AI surfaces describe a business before that date was produced by a different generator than the one answering today.
Primary source: Google (Gemini models launch post, 21 July 2026)
What Google actually said
The primary source is Google's launch post of 21 July 2026. Three models shipped: Gemini 3.6 Flash (the workhorse tier — Google says it "reduces output token usage by 17% compared to 3.5 Flash," priced at $1.50/1M input and $7.50/1M output), Gemini 3.5 Flash-Lite ("the fastest model in the 3.5 series… as measured by Artificial Analysis, it runs at 350 output tokens/s," $0.30/1M input and $2.50/1M output, "designed for both low-latency tasks and tasks where high throughput is critical… like agentic search and document processing"), and Gemini 3.5 Flash Cyber, a limited-access security model. The Gemini API changelog records both `gemini-3.6-flash` and `gemini-3.5-flash-lite` reaching general availability the same day, and adds a detail with direct consequences for measurement: the sampling parameters `temperature`, `top_p` and `top_k` are now deprecated on the latest models.
The inference we are not making
Google wrote "rolling out in Google Search." It did not write AI Overviews, and it did not write AI Mode. It has been that specific before when it wanted to be: at I/O on 19 May 2026, Search VP Elizabeth Reid announced Gemini 3.5 Flash — the full Flash tier, not Flash-Lite — "as the new default model in AI Mode for everyone globally." This week's post names no surface. Those are the two surfaces every AI-visibility conversation is actually about, and it is a short, natural leap to assume a search-deployed model is the one generating them — but Google has not said so, and we are not going to publish the leap as a fact. Take the sentence at exactly its stated width: a cheap high-throughput model is now serving some part of Search, and Google is not itemising which part.
What a generator swap changes, and what it doesn't
AI answers are produced in two layers that fail independently. The retrieval layer decides which documents the engine pulls; the generation layer decides what gets said about them, which businesses get named, and in what order. A model swap replaces the second layer while leaving the first largely alone — which is why the panic response (rewrite the site) is usually the wrong one and the boring response (re-baseline, then keep working the corpus) is usually right.
Our own corpus says where the durable work sits. Across 511 ChatGPT web-search answers, covering 172 audits of local businesses, answers averaged 5.5 distinct sources and the business's own website appeared in 28% of them — the rest was review aggregators, "best-of" directories and third-party write-ups. Those documents do not change when a vendor swaps a model. Who gets named out of them can change overnight, without an announcement, and without anything on your site being different.
Three things to do about it this week
- 1Date-stamp the baseline and refuse to diff across the boundary. A Google-surface reading from 10 July and one from today are two experiments, not two points on a trend line. If you report AI visibility to a client monthly, the July report needs a line saying the generator changed mid-month — otherwise you will explain a model swap as the effect of your own work, in whichever direction it moved.
- 2Do not rewrite the site because a model shipped. The evidence base the engine reads is third-party and slow-moving; the generator is neither. Work that survives a swap is work that adds documents to the corpus — directory records, review depth, cited third-party coverage — not a schema tweak deployed the day after a launch post.
- 3If you measure Gemini through the API, know what your proxy is now worth. Every AI-visibility tool we know of — ours included; we call `gemini-2.5-flash` with grounding — measures a model, not the consumer Search surface. That gap just widened, and with `temperature`, `top_p` and `top_k` deprecated on the newest models you can no longer pin sampling to make runs comparable. The remaining honest defence is repeated sampling: ask the same question several times and report how often the business appeared, not whether it appeared once.
For an agency, the practical version of all three is a dated re-baseline across the client roster rather than a scramble on the loudest account — which is what the $249 Agency 5-Pack is: five white-label audits, same method, same week, so a portfolio has one comparable before-and-after line instead of five anecdotes. For a single business, the free AI visibility check re-runs the buyer's question across engines and reports what came back today.
See your number
A free 60-second check shows what AI says about you.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.