A content strategy for AI Mode is a plan for covering the hidden sub-queries Google generates from a single prompt, using enough separately answerable pages and sections that your domain can fill several of those branches at once. It is not a longer-content strategy and not a more-pages strategy — both fail for the same reason.
We tracked 19,412 AI Mode citations across 60 prompts over eight weeks to settle the argument. Neither camp won. The variable that predicted citations had almost nothing to do with word count.
This matters because the two camps give opposite instructions with real cost attached. Consolidate everything into definitive guides, says one. Publish a page per sub-question, says the other. Until now the evidence on either side has been anecdotal.
What the query fan-out does to a page
Query fan-out is the process where an AI search system converts one user prompt into several concurrent sub-queries, retrieves results for each, then synthesises a single answer. Google describes it in its guidance on AI features in Search as issuing "a set of related searches" across its index to assemble one response.
Your page is no longer competing for one query. It is competing separately for each branch of a tree it never sees. In our sample the median parent prompt produced 9.3 distinguishable facets, ranging from 4 to 21.
That creates more citation slots than a classic SERP has ranking positions — and it changes what "covering a topic" means. A page can be excellent on the parent question and invisible to seven of the nine branches. For the retrieval mechanics behind those branches, see our breakdown of how one prompt becomes dozens of hidden searches.

Why existing advice doesn't answer the question
Guides on depth versus breadth converge on the same conclusion: build topic clusters, balance both, link them semantically. Not wrong. Just untested at the level where the decision gets made — the individual URL.
None of them publish citation counts per URL. Meanwhile the strongest public datasets measure URL attributes rather than content architecture. Otterly's study of 1,028,959 cited URLs found path depth correlated with citation volume at r = +0.002 and URL length at r = −0.025 — effectively zero. The shape of the address tells you nothing.
So the honest state of the evidence was: URL formatting doesn't matter, clusters are a good idea in the abstract, and nobody had measured whether one deep page or a set of narrow pages captures more of the fan-out.
How we tested depth against breadth
We tracked 60 parent prompts daily in Google AI Mode for eight weeks, logged every cited URL, and isolated the 41 cases where one domain ran both a comprehensive hub and three or more narrow pages on the same topic. Same domain, same topic, same authority, different architecture — that matched-pair design is what makes the comparison usable.
- Prompts: 60 commercial-investigational and informational prompts across B2B SaaS, data infrastructure, HR tech and cybersecurity.
- Window: 30 March – 24 May 2026, pulled daily, US locale, desktop, logged-out sessions.
- Volume: 3,360 scheduled pulls, 2,904 usable AI Mode responses, 19,412 citation instances across 6,730 unique URLs.
- Deduplication: one credit per URL per response;
google.comself-citations excluded. Counting rules follow our guide to counting AI citations without inflating the result. - Facet reconstruction: we clustered the answer sentences each citation supported, then labelled the distinct sub-question each cluster resolved. That gives an observed facet list per prompt rather than a guessed one.
- Matched sets: 41 domains where a hub (≥2,500 words, ≥6 facets addressed) coexisted with ≥3 narrow pages (<1,500 words, 1–2 facets) on the same topic.
Limits worth stating. Facet labelling is human judgement, not a model output — two analysts agreed on 87% of labels and reconciled the rest. Logged-out US desktop sessions do not capture personalisation. And 41 matched sets is enough to separate a 3× effect from noise, not enough to rank verticals against each other.
Mean citations per URL across the whole set was 2.9, median 2 — higher than Otterly's 1.9 mean, as expected, since our window is eight weeks of repeated daily prompts rather than a single snapshot.
What the data showed: depth wins per URL, breadth wins coverage
Comprehensive hubs earned a median of 5.4 citations each over eight weeks across a median of 3 distinct facets. Narrow pages earned 1.6 citations across 1 facet. But the typical spoke set out-earned the single hub on total volume — 7.1 citations to 5.4.
| Metric (median, 8 weeks) | Comprehensive hub | Single narrow page | Full spoke set (3–6 pages) |
|---|---|---|---|
| Citations per URL | 5.4 | 1.6 | 1.6 |
| Total citations | 5.4 | 1.6 | 7.1 |
| Distinct facets cited in | 3 | 1 | 4 |
| Facets covered but never cited | 4 | 0 | 1 |
Two things follow. A deep page is more efficient per unit of work — one URL doing the job of three. But efficiency is not the goal; coverage is. Spoke sets won on total citations by roughly 31% and on facet spread by a full point.
The cannibalisation fear did not show up. Where hub and spokes both existed, combined citations reached a median of 12.9, and only 6 of 41 hubs lost volume after spokes were published.
The variable that actually predicted citations
Word count did not predict how many fan-out facets a page got cited in. The number of sections opening with a direct, self-contained answer did — r = 0.61 against r = 0.08 for length.
Of the 41 hubs, 14 (34%) earned citations across four or more distinct facets and averaged 11.2 citations. The other 27 were cited in two facets or fewer and averaged 2.8 — statistically indistinguishable from a single narrow page.
The two groups were nearly identical in length. Multi-facet hubs averaged 3,410 words; single-facet hubs 3,180. A 230-word gap is noise.
What separated them was internal shape. Multi-facet hubs had a median of 9 H2 or H3 sections opening with a direct answer to a question a person would type on its own. Single-facet hubs had 2. Everything else was narrative connective tissue — readable, not retrievable as a standalone passage.
So "comprehensive" is not a word count. A comprehensive page is one carrying many independently answerable blocks. A 4,000-word essay with three headings is a long narrow page wearing a costume.
What a direct-answer section looks like
The test is mechanical: cut the first 40–60 words out of the section, show them to someone with no other context, and see whether they answer a question that person could have typed alone.
- Retrievable: "Query fan-out is the process where an AI search system converts one user prompt into several concurrent sub-queries, retrieves results for each, then synthesises a single answer."
- Not retrievable: "Now that we've covered the basics, let's look at how this plays out in practice — and why so many teams get it wrong."
Second version may read better in sequence. It cannot be lifted. Same page, same word count, different citation ceiling. This matches the format patterns in which content formats AI search cites most, where structural answerability outperforms length in every vertical we track.

The blind spot breadth covers and depth cannot
Across our 60 prompts, 23% of observed facets were not addressed anywhere on the matched hub page. Those branches went to spokes, competitors, or third-party sources — forums, comparison sites, vendor documentation.
This is the honest case for breadth, and it has nothing to do with keyword coverage. It is a forecasting problem. When you plan a hub, you write the sub-questions you can anticipate. The fan-out decomposes intent its own way, which regularly surfaces angles no content brief would have listed: pricing objections, migration friction, regulatory edge cases, "does this work with X."
You cannot pre-empt those inside one page without bloating it into something nobody reads. You catch them with focused pages published in response to observed facets — which requires seeing the facets first. That reverses the usual planning order: measure the fan-out, then commission the content.
Community threads absorb a large share of the unpredicted branches, which is why how Reddit threads become AI recommendations belongs in the same plan as your own publishing schedule.
The three-facet rule for splitting or consolidating
Give a subtopic its own URL when it needs three or more distinct answerable facets. Keep it as a section of the hub when it needs one or two. In our matched sets, subtopics split below that threshold produced pages averaging 0.9 citations — barely above zero — while subtopics kept as hub sections above it left facets uncovered.
Run each candidate subtopic through five checks:
| Signal | Keep as a hub section | Split into its own URL |
|---|---|---|
| Distinct answerable facets | 1–2 | 3 or more |
| Can you write 150+ unique words per facet? | No | Yes |
| Do you hold first-hand data or examples for it? | No | Yes |
| Does it carry its own buying or decision question? | No | Yes |
| Would the page stand up if the hub disappeared? | No | Yes |
Three or more "split" answers means split. Fewer means it is a section, and forcing it into a URL produces a page that competes with your hub and loses.
Two cases where the rule pointed in opposite directions
Both ran on the same tracking setup: daily AI Mode pulls, citations attributed per URL, eight-week window after the change shipped.
Consolidation: a payroll SaaS with 14 thin pages
14 pages averaging 640 words, one per payroll sub-question. Under the three-facet test, ten scored 1–2 facets. Those ten merged into a single hub of 3,800 words with 11 direct-answer sections; the four scoring 3+ stayed standalone and were rewritten. Merged URLs were 301-redirected into the hub.
| Before | After 8 weeks | |
|---|---|---|
| Citations per week | 4 | 21 |
| URLs cited | 3 | 5 |
| Facets covered | 2 | 9 |
Nothing moved for the first five weeks. Recrawl and re-embedding lag is real, and teams judging at week two will conclude the change failed.
Splitting: a data-infrastructure vendor with one 7,400-word guide
The opposite starting point. One enormous guide, four direct-answer sections, cited in 3 facets at 9 citations per week. Six subtopics inside it each cleared the three-facet threshold, so they were extracted into their own pages and the hub was rewritten down to 2,600 words with clean handoffs to each spoke.
After eight weeks: 26 citations per week across 14 facets. The hub itself dropped from 9 to 6 citations per week. All net gain came from the spokes — anyone measuring only the original URL would have recorded a 33% loss and reverted a change that nearly tripled topic-level citations.

Why AI Mode rewards domains that can supply several URLs
AI Mode does not behave like a system looking for one best page. SE Ranking's analysis of AI Mode responses found roughly 12.6 links per answer drawn from only about 2.29 unique domains — about 5.2 URLs from the same domain per response.
Read that alongside the matched-pair results and the picture sharpens. AI Mode concentrates trust on few domains, then pulls multiple pages from each to fill the fan-out branches. A domain with one great page can supply one or two slots. A domain with a hub and five genuinely distinct spokes can supply six.
That is topical authority stated concretely: authority converts into citation volume only if you have enough distinct, answerable URLs for the model to draw on. Which is also why the domain-selection mechanics in how ChatGPT, Perplexity and Gemini decide which brands to cite sit upstream of everything here — you have to clear the domain filter before URL count does anything for you.
What to publish when you have no data of your own
The three-facet rule assumes you know your facets. Most teams do not yet. Three page types earn facets fastest because they answer sub-questions that generic content cannot:
- Original statistics. A number nobody else has is the one asset that gets cited by name rather than paraphrased. Small brands with a survey of 300 customers can out-cite incumbents on a facet — see original statistics as a citation magnet.
- Comparison and criteria pages. "X vs Y", "how to choose", "what to look for" facets appear in nearly every commercial fan-out we track and are frequently answered by third parties instead of vendors.
- Objection and edge-case pages. Pricing pushback, migration cost, compliance limits, integration gaps. These are the facets hubs skip and the ones with the least competition.
One caution on the third type: publishing about the topic is not the same as publishing on it. If your facts live on Medium or LinkedIn as well as your own domain, who wins the citation when content lives in two places determines whether the work accrues to you.
Where a breadth strategy quietly fails
Publishing a page per sub-query at scale is the one approach Google names as a policy risk. Its guidance is explicit: creating many pages primarily to manipulate rankings, with little value per page, falls under the scaled content abuse spam policy.
The three-facet rule is a practical guard. A subtopic clearing three distinct answerable facets, carrying its own decision question and able to stand without the hub is a real page. A subtopic clearing none of those is a doorway with better manners.
The failure mode in audits is rarely deliberate spam. It is a content calendar built from a keyword export, where 40 near-identical pages get commissioned because 40 rows existed in a spreadsheet. Looks like breadth, behaves like thin content. In our sample, pages under 800 words with a single direct-answer section averaged 0.4 citations across eight weeks.
Auditing your own pages against the fan-out
A working content strategy for AI Mode starts with observation, not an outline. Seven steps:
- List the parent prompts your buyers actually use, including multi-turn follow-ups, not just head keywords.
- Capture the facets. Log AI Mode answers for each prompt and cluster the answer sentences by the sub-question they resolve. That is your real fan-out map.
- Map facets to URLs. For each facet, record which of your pages — if any — is cited, and which competitor or third-party source is.
- Count direct-answer sections per URL. Any heading whose first 40–60 words do not resolve a standalone question is not a retrievable block.
- Apply the three-facet rule to every subtopic to decide split, keep or merge.
- Rebuild internal links so the hub points to each spoke with descriptive anchors and each spoke points back. That is how you signal the cluster relationship.
- Re-measure over six to eight weeks. Anything shorter is noise.
Step 4 is where most audits find the problem. Teams expect to discover they need more pages and instead discover their existing pages have two retrievable blocks each.
Steps 2 and 3 are the manual bottleneck. If you are doing this at more than a handful of prompts, compare what the AI Overviews and AI Mode tracking tools actually capture — most report mentions, far fewer expose the cited URL per answer, which is the field this entire audit depends on.
Measuring the change without fooling yourself
Track citations per URL and facets covered per topic — not total mentions. Total mentions rise when the fan-out widens, which happens for reasons unrelated to anything you shipped.
Three measurement rules from running these tests:
- Attribute at the URL level, then roll up to topic. The splitting case above looked like a 33% loss at the hub and a 189% gain at topic level. Same change.
- Hold a control. Leave a comparable topic untouched over the same window. Without it you cannot separate your edit from a model update — AI Mode answers shift under you regardless of what you publish.
- Watch facet coverage, not rank. There is no rank in AI Mode. The closest thing to a position is the count of distinct sub-questions where you appear.
The same discipline transfers across engines. A page structured for facet coverage tends to lift share of voice in ChatGPT, Perplexity and Copilot too, because all of them decompose prompts and retrieve passages. But facet lists differ per engine — track them separately rather than averaging into one score.
The short version
Depth and breadth were never the real variables. Answerable facets per URL is the variable, and it explains both camps at once: deep pages win when they carry many independently readable blocks, and narrow pages win when they cover facets the hub genuinely cannot.
One number to take away: word count correlated with facet coverage at r = 0.08. Direct-answer section count correlated at r = 0.61. Stop writing longer. Start writing more separately answerable.
Frequently asked questions
What is a content strategy for AI Mode?
It is a plan for covering the hidden sub-queries Google generates from one prompt, with enough separately answerable pages and sections that your domain can fill several branches of the fan-out at once. It replaces the keyword-per-page model with a facet-per-block model, and it starts by measuring the fan-out rather than guessing it.
Should I consolidate my thin pages or publish more of them?
Apply the three-facet rule to each page. Pages covering one or two facets should be merged into a hub — in our sample they averaged 1.6 citations regardless of how many existed. Pages covering three or more distinct facets, with their own decision question, should stay and be deepened.
Does a longer page get cited more often in AI Mode?
No. In our 41 matched sets, multi-facet hubs averaged 3,410 words and single-facet hubs 3,180 — a difference within noise. Length helps only when the extra words create additional self-contained sections answering distinct questions.
How many sub-queries does AI Mode generate per prompt?
Our tracked prompts produced a median of 9.3 distinguishable facets, range 4 to 21. Complex commercial prompts fan out considerably wider than informational ones.
Will publishing hub and spoke pages on the same topic cannibalise my citations?
Rarely. Combined hub-plus-spoke sets earned a median of 12.9 citations against 5.4 for the hub alone, and only 6 of 41 hubs lost volume after spokes launched. SE Ranking's finding of roughly 5.2 URLs cited per domain per response points the same way — AI Mode routinely cites several pages from one site.
How long before an architecture change shows up in AI citations?
Budget six to eight weeks. In both cases above, nothing moved for the first five weeks while pages were recrawled and re-embedded. Judging results at week two makes a working change look like a failure.
Do I need a hub at all, or just spokes?
You need the hub, but for a different reason than usual. Spoke sets out-earned hubs on total citations, yet the hub is what holds the cluster together for internal linking and covers the parent question itself — which is its own facet in every prompt we tracked. Publishing spokes with no hub leaves the broadest sub-query to a competitor.