AI search news · ·

A title-shape A/B on 22 posts read inconclusive on its pre-registered date: question titles were cited by 0.5 of 4 AI engines per post, statement titles by 0.4, and 20 of 20 posts scored 2 or fewer

On 2026-09-13, the read date written down on 2026-08-23 before the first post was enrolled, the question-shape A/B on this site's daily posts scored 0.5 cited engines of 4 per post for question titles and 0.4 for statement titles, and the instrument's verdict line reads "INCONCLUSIVE — reported, not iterated". 22 posts were enrolled between 2026-08-24 and 2026-09-11, 11 in each arm, the arm fixed by the enrollment sequence before the headline was written. 20 had at least one Sunday measurement and 11 had the second-Sunday measurement the rule scores: 6 question posts, 5 statement posts, against a bar of 4 each. The 0.1-cell difference sits inside the band the rule names inconclusive, and the excess is on one engine, Gemini, where one question post of 6 was cited and 0 statement posts of 5. Beside the verdict: 20 of 20 measured posts scored 2 or fewer engines of 4 on every Sunday, 14 of 20 were never cited by any engine, and no post was cited by Perplexity or Claude in any scored cell. Enrollment closed with the read; title shape is editorial judgement per story from here.

Method, stated before the numbers

The unit is one ordinary daily post on this site. The variable is the shape of its headline and URL: a question (the h1 begins with an interrogative and carries a question mark; the slug begins with an interrogative token) or a statement (a number-led sentence; neither the h1 nor the slug carries either mark). The arm was assigned by the enrollment sequence, question for even sequence numbers and statement for odd, and the newsroom learned its arm before writing the h1 or slug. After each post shipped, its URL was bound to its cell; the server refuses an h1 or slug whose shape contradicts the arm. The Monday standing post was exempt. Every post was enrolled under the verbatim user question it answers.

The instrument is the same panel that runs our weekly citation read: every Sunday at 12:37 UTC each enrolled post's question was put to ChatGPT, Perplexity, Gemini and Claude, three samples each, through the API. A cell is one engine on one post; it counts as cited when any of its three answers cites askedabout.com. A post is measured on every Sunday that falls at least 3 days after its URL was attached, and the rule scores the second such Sunday, so both arms are read at the same age. The primary metric is cited cells of 4 per post, averaged per arm.

The rule, written 2026-08-23 before any enrolled post existed: the read happens on the first Sunday with at least 4 posts per arm carrying a second-Sunday measurement, with one extension allowed if either arm is short. Question mean at or below statement mean gives the question style no credit. Question mean above statement mean by at least 0.5 cells per post, with the excess on at least 2 engines, banks the style. Anything between is inconclusive, reported and not iterated. The read is code, committed the same night, and the 2026-09-13 numbers below are its output, reproduced field for field by the committed script `content/news/ab_read_cut.mjs`, which writes `content/news/ab-read-cut-2026-09-13.json` and passes 18 of 18 self-checks, among them that the scored measurement is never a post's first Sunday and that the arm follows sequence parity on all 22 cells.

The read, 2026-09-13

ArmEnrolledMeasuredScored (2nd Sunday)Mean cited engines of 4Posts with a cited cellChatGPTGeminiPerplexityClaude
Question (Q)111060.52 of 62 of 61 of 60 of 60 of 6
Statement (S)111050.42 of 52 of 50 of 50 of 50 of 5

Both arms clear the bar of 4, so the read is taken on its date and not extended. The difference is 0.1 cells per post, above 0 and under 0.5. The question arm's excess appears on one engine, Gemini, where one question post was cited and no statement post was; on ChatGPT the arms are 2 of 6 and 2 of 5; on Perplexity and Claude both arms are 0. The verdict line printed by the instrument is "INCONCLUSIVE — reported, not iterated". By story kind, the covariate the pre-registration named: first-party posts Q 0.6 (5 posts) against S 0.4 (5 posts); the one scored third-party post is a question post at 0, with no statement post to set against it.

Every enrolled post

Sequence, arm, kind, ship date, cited engines of 4 on each Sunday it was measured, and the scored cell. A dash is a Sunday before the post was 3 days old; posts 11 to 19 have one measurement and await their second Sunday, which the rule does not read; posts 20 and 21 shipped too late for any measurement by 2026-09-13.

SeqArmKindShippedAug 30Sep 6Sep 13ScoredEngines in the scored cell
0Qfirst-party2026-08-240000none
1Sfirst-party2026-08-250000none
2Qfirst-party2026-08-260000none
3Sfirst-party2026-08-27-011ChatGPT
4Qfirst-party2026-08-28-011ChatGPT
5Sfirst-party2026-08-29-000none
6Qfirst-party2026-08-30-122ChatGPT, Gemini
7Sfirst-party2026-08-31-000none
8Qthird-party2026-08-31-000none
9Sfirst-party2026-09-01-211ChatGPT
10Qfirst-party2026-09-02-000none
11Sfirst-party2026-09-03--0awaiting-
12Qthird-party2026-09-03--0awaiting-
13Sfirst-party2026-09-04--0awaiting-
14Qthird-party2026-09-04--0awaiting-
15Sfirst-party2026-09-05--0awaiting-
16Qfirst-party2026-09-06--0awaiting-
17Sfirst-party2026-09-07--2awaiting-
18Qfirst-party2026-09-08--0awaiting-
19Sfirst-party2026-09-09--1awaiting-
20Qfirst-party2026-09-10---not measured-
21Sfirst-party2026-09-11---not measured-

What the pre-registration bought

Without it, the claim on the table would have been made from one or two posts. Post 6, a question post, was cited by 2 engines of 4 at its second Sunday, the best score on the record; read alone against its two neighbours at 0, it argues for question titles. Post 9, a statement post, was cited by 2 engines on its first Sunday and 1 on its second; read alone it argues the other way. The rule fixed on 2026-08-23 named the metric, the age at which each post is read, the bar per arm, the read date, and three outcomes with their thresholds, so the 2026-09-13 numbers could only be reported against them. The house style that ran for three weeks earned no credit and no penalty, and the newsroom is not extending the test a second time, because the rule said it would not.

What it did not test

The population beside the verdict

FactCount
Measured posts scoring 2 or fewer engines of 4, on every Sunday20 of 20
Measured posts cited by at least one engine on at least one Sunday6 of 20
Measured posts never cited by any engine on any Sunday14 of 20
Engine-cells citing askedabout.com in the 2026-09-13 run8 of 80
Samples citing askedabout.com in the 2026-09-13 run12 of 240
Scored Perplexity and Claude cells citing askedabout.com0 of 22

The arm difference is small against this floor. Whatever the title shape, a new post on a site of this size is cited by at most 2 of the 4 engines when asked its own question two weeks after it ships, and by none in 14 cases of 20. On ChatGPT 4 of the 11 scored posts were cited; on Gemini 1; on Perplexity and Claude none. The question of which engine ever names a given page, asked about a business rather than a blog post, is what the free check reads, and the page on showing up in ChatGPT recommendations sets out what the engines cite when they do.

What changes

Enrollment closed on 2026-09-13 with the read. Title and slug shape return to editorial judgement per story. The 22 attached cells stay in the panel and are measured on Sundays as history; the rule reads once and the nine posts awaiting a second Sunday are not scored. The next content variable under a read on this site is not a title, it is a link: from today, three posts Googlebot re-fetches carry a dated link to each of the six posts it has not fetched, read on 2026-09-20.

Why was the test not extended when the result was inconclusive

The pre-registration allowed one extension only when an arm was short of 4 scored posts. Both arms cleared the bar on 2026-09-13, so the read was taken. Extending after seeing an inconclusive result is the reinterpretation the rule was written to prevent.

Why is the second Sunday scored and not the first

Posts ship on every weekday, so the first Sunday at least 3 days after a post is attached falls at ages 4 to 10 days on this record. The second falls at ages 11 to 17, and both arms are read there so that age does not favour either arm.

Does 0.5 cited engines of 4 mean question titles work

No. The difference to the statement arm is 0.1 cells per post, the rule's inconclusive band is 0 to 0.5, and the excess is on one engine, one post. The population fact is the stronger reading: 20 of 20 posts scored 2 or fewer of 4 whatever their title.

See your number

See which businesses AI names when your client's buyers ask.

Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.

Check a client's AI visibility

Begin your check

Free · 60 sec

No account · No card · 3 buyer questions, 2 engines

By running a check you agree to our Terms and Privacy Policy.

Who runs this