Ask Google the same question twice — once on the classic results page, once in the conversational tab — and you get two different bibliographies. In our tracking panel, AI Mode vs AI Overviews shared only 18.6% of their cited URLs across 731 queries where both surfaces responded. The answers usually agreed. The sources rarely did.
That gap is the single most misread fact in Google AI visibility work right now. Teams celebrate an AI Overviews citation, assume the AI Mode citation came with it, and report one number upward. Roughly four times out of five, it didn't come with it.
This piece covers what we measured, why the two surfaces pull from different pools, and — the part nobody else has published — what separates pages that win both from pages that win only one.

What is the difference between AI Mode and AI Overviews?
AI Overviews is a synthesized answer block that appears above classic Google results on some queries. AI Mode is a separate conversational tab that always responds, runs many hidden sub-searches per prompt, and supports follow-up turns. Overviews summarize; AI Mode researches.
| AI Overviews | AI Mode | |
|---|---|---|
| Where it lives | Block above classic results | Its own tab |
| Fires on every query? | No — 60.9% of our prompts | Yes — 100% |
| Retrieval style | Extraction from ranked results | Fan-out into many sub-searches |
| Follow-up turns | No | Yes |
| Typical citation count | ~6 URLs | ~13 URLs |
| Cites the brand's own site | 8.7% of citations | 21.3% of citations |
| Rewards | One clean answer, early | Many adjacent sub-answers |
That structural difference drives everything else. AI Overviews must fit a compact block above a results page it doesn't replace, so it cites a short, curated list. AI Mode builds a longer answer from many parallel retrievals, so it cites a wider and less predictable set.
Google's own documentation treats them as one family. Its Search Central guidance on AI features states there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary," and that both surfaces "display a wider and more diverse set of helpful links associated with the response than with a classic web search."
Accurate as policy. Not much help operationally — because in practice the two surfaces disagree about which links are the helpful ones.
How we ran the study: 1,200 prompts, two surfaces, six weeks
We built a prompt library of 1,200 queries drawn from the tracked topic sets of 40 B2B SaaS and technology brands — the real questions their buyers ask, not head terms. Every prompt ran daily through both Google AI Mode and Google AI Overviews between April 6 and May 17, 2026, from a US locale, desktop, logged out, with no personalization carried between runs.
For each response we captured every outbound URL visible in the initial answer and its attached source panel. We did not expand nested link cards or issue follow-up turns — that keeps the comparison honest, since AI Mode's expandable panels can inflate a citation count several times over without the user ever seeing those links.
Total collected: 102,417 citation instances, deduplicated to URL and root-domain level per query per day.
Analysis ran on the paired set — the 731 queries where both surfaces produced a response on the same day. Comparing a surface that fired against one that didn't would manufacture a divergence that isn't there.
Limitations, stated up front: US English, desktop, B2B-weighted, six-week window, single tracking panel. Consumer, mobile, and non-English behaviour differ — sometimes sharply, as we've documented in multilingual AEO testing. Numbers below are one panel's observations, not a Google-published figure.
The headline number: 18.6% of cited URLs appear on both surfaces
Across the paired set, the median AI Mode response cited 12.6 unique URLs. The median AI Overviews response cited 5.9. They shared a median of 2.9.
| Metric (paired set, n=731) | AI Mode | AI Overviews |
|---|---|---|
| Responded to prompt | 1,200 / 1,200 (100%) | 731 / 1,200 (60.9%) |
| Median unique URLs cited | 12.6 | 5.9 |
| Median unique root domains | 9.1 | 5.2 |
| Responses with zero outbound citation | 1.8% | 7.4% |
| Citations pointing to the brand's own domain | 21.3% | 8.7% |
| Citations to third-party editorial or review sites | 34.6% | 47.9% |
Overlap depends heavily on how you measure it:
| Comparison level | Overlap |
|---|---|
| Exact URL (Jaccard) | 18.6% |
| Root domain (Jaccard) | 39.8% |
| Top 3 cited URLs only | 14.1% |
| Brand named in prose, linked or not | 57.2% |
That last row matters more than it looks. Brands get named on both surfaces roughly three times as often as they get linked on both. If your reporting counts only clickable citations, you're undercounting real presence — and if it counts only mentions, you're overcounting your ability to send traffic. Any serious AI share-of-voice model needs both columns.
Our 18.6% URL overlap runs meaningfully higher than the 13.7% Ahrefs reported from 540,000 query pairs on general-population queries. The likely explanation is source-pool depth: a query like "best running shoes" has thousands of credible candidate pages, so two retrieval systems can diverge almost completely. "Best reverse ETL tool for a Snowflake stack" has maybe forty. Thin pools force agreement.
The practical read: the more niche your category, the more your two surfaces will look alike — and the more a single lost citation costs you on both at once.
Winning AI Overviews predicts AI Mode twice as well as the reverse
The overlap is not symmetric, and this is the finding we'd hand a marketer first.
| If a page is cited in… | …probability it is also cited in the other surface |
|---|---|
| AI Overviews | 49.2% |
| AI Mode | 23.0% |
An AI Overviews citation is roughly 2.1× more predictive of an AI Mode citation than the reverse. Half of everything Overviews picks also shows up in AI Mode. Fewer than a quarter of AI Mode's picks make it into Overviews.
Read it as filter strength. AI Overviews selects from a narrower, higher-bar shortlist that AI Mode's broader net tends to also catch. AI Mode's long tail — the pages it surfaces at citation position 8 through 13 — mostly never clears the Overviews bar.
Two consequences for planning:
- If you already win AI Overviews for a query, don't spend budget chasing AI Mode on that same query. You probably have it. Verify, then move on.
- If you win AI Mode only, treat it as a partial win, not a leading indicator. Three times out of four it will not convert into an Overviews citation on its own.
This asymmetry also explains a reporting pattern we see constantly: an agency shows a client an impressive AI Mode citation count, the client checks a live search on their phone, sees no mention in the Overview block, and trust evaporates. Both numbers were correct.
Overlap collapses exactly where commercial intent lives
Aggregate overlap hides the useful signal. Segmented by query type, the spread runs from 34% down to 9%.
| Query type | Example prompt | URL overlap | AI Overviews trigger rate |
|---|---|---|---|
| Definitional | "what is answer engine optimization" | 34.1% | 88% |
| How-to / implementation | "how to add schema to a SaaS pricing page" | 24.6% | 71% |
| Head-to-head comparison | "Segment vs RudderStack" | 21.7% | 64% |
| Brand evaluation | "is Clay any good for outbound" | 15.2% | 39% |
| Pricing / cost | "how much does an AI visibility tool cost" | 11.8% | 47% |
| Shortlist / best-of | "best AI search monitoring tools for agencies" | 9.3% | 41% |
Definitional queries converge. Purchase-adjacent queries diverge. On "what is X" prompts, both surfaces reach for the same encyclopedic and canonical explainers — which is precisely why well-built definition and glossary pages remain the cheapest double win available.
On "best X for Y" prompts, overlap falls to 9.3%. Those queries have no canonical answer, so AI Mode's fan-out assembles a shortlist from a dozen scattered listicles, forum threads and vendor pages, while AI Overviews — when it fires at all, which is 41% of the time — picks the two or three it trusts most.
That single row reframes the strategy. The queries with the highest commercial value have the lowest source overlap, which means they need two separate optimization approaches, not one. A brand that treats "get cited in the Overview block" and "get recommended in the AI Mode tab" as one project will underperform on both.
Why the two surfaces diverge: fan-out, extraction, and trigger thresholds
Three mechanisms explain nearly all of the gap.
AI Mode fans out; AI Overviews extracts
AI Mode decomposes one prompt into many hidden sub-searches, retrieves for each, then synthesizes. A prompt about "best AI search monitoring tools for agencies" quietly becomes searches for agency reporting features, white-label options, per-client pricing, platform coverage and so on. Each sub-search contributes its own sources. We cover the mechanics in how one prompt becomes dozens of hidden searches.
AI Overviews behaves much more like a compressed extraction layer sitting on top of the classic index — closer in spirit to the old featured snippet than to a research agent, a lineage we traced in what changed for the content that used to win position zero.
Fan-out rewards pages that answer many adjacent sub-questions. Extraction rewards pages that answer one question cleanly and early. Different page shapes — which is why so few pages are good at both.
AI Overviews often doesn't fire at all
In our paired-set analysis, AI Overviews responded to 60.9% of prompts while AI Mode responded to 100%. Third-party testing reports a wider gap: Otterly ran 100 top German queries and found AI Mode triggering on every query while AI Overviews appeared for 49.
Our higher rate is almost certainly the B2B, information-seeking skew of our library — Overviews suppresses more aggressively on ambiguous consumer and YMYL-adjacent terms.
Either way, the operational point holds: on roughly 40% of your tracked prompts, AI Overviews is not a channel at all. Any dashboard that averages the two surfaces into a single "Google AI visibility" score is averaging across a surface that wasn't there.
The two surfaces weight owned content differently
AI Mode sent 21.3% of its citations to the queried brand's own domain. AI Overviews sent 8.7% — under half. Overviews leans on third-party validation; AI Mode is far more willing to quote you about yourself.
Practical consequence: your own site is a viable AI Mode lever and a weak AI Overviews lever. Rewriting your pricing page can move AI Mode within a crawl cycle. Moving Overviews on the same query usually requires a third-party mention you don't control. This shows up most starkly on brand-evaluation prompts, where the ability to influence your own narrative diverges sharply between surfaces — the same dynamic that makes employer-brand questions behave so differently from product questions.
What double winners have in common
We pulled the three cohorts apart and profiled them: 1,412 URLs cited by both surfaces, 4,720 cited by AI Mode only, 1,455 cited by AI Overviews only.
| Signal | Double winners (n=1,412) | AI Mode only (n=4,720) | AI Overviews only (n=1,455) |
|---|---|---|---|
| Median word count | 1,850 | 2,940 | 1,120 |
| Distinct questions answered on page | 9.4 | 7.8 | 4.2 |
| Answer density (answers per 1,000 words) | 5.1 | 2.7 | 3.8 |
| Direct answer within first 200 words | 81% | 44% | 73% |
| Contains a comparison table or spec list | 57% | 24% | 31% |
| Cited by ≥3 independent domains on topic | 68% | 33% | 61% |
| Page sits on the brand's own domain | 34% | 61% | 22% |
Double winners are not the longest pages. They're the densest — 5.1 discrete answers per thousand words, versus 2.7 for the AI Mode-only cohort.
We counted a "distinct question answered" as a self-contained passage that resolves one question without requiring the reader to have read the section above it — the same unit an extraction layer can lift. Two coders scored a 200-page sample; disagreements were reconciled by re-reading against that rule.
AI Mode-only pages are long-form narrative: 2,940 median words covering 7.8 subtopics, with the lead answer buried past word 200 more than half the time. Fan-out finds them because they touch many sub-questions somewhere in the body. Overviews skips them because there's no clean block to lift.
AI Overviews-only pages are the mirror image: 1,120 words, crisp lead answer, 4.2 subtopics. Extraction loves them. Fan-out runs out of surface area after two or three sub-searches and moves on.
The double-winner shape is a tight, quotable answer in the first 200 words, followed by nine or ten more self-contained answers, each independently liftable. Structure, not length.
One line worth reading twice: 68% of double winners were cited by three or more independent domains on the same topic, versus 33% of AI Mode-only pages. Off-site corroboration is the strongest non-structural predictor of winning both. You cannot write your way to it — which is also why AI Mode-only wins are the ones a competitor can take from you fastest.

Diagnose any page in one pass: the four quadrants
Every tracked page falls into one of four states per query. Each has a different fix.
- Wins both. Nothing to do. Protect it — log its structure as your internal template and check monthly for source-set churn.
- Wins AI Mode only (the largest cohort, 62% of cited URLs). Breadth, no extractable lead. Fix: rewrite the opening 150 words into a direct, self-contained answer to the exact query; add one comparison table; leave the body alone.
- Wins AI Overviews only (19%). Answers one question well and stops. Fix: extend with 5–8 additional self-contained sub-answers under descriptive H2s, mapped to the fan-out questions your monitoring surfaces. Do not pad the existing section.
- Wins neither. Before touching the page, check whether any of your domain's pages win the query. If a competitor's page wins both, study its answer density before rewriting anything. If nobody wins both, that's an open slot — the cheapest citation you'll ever earn.
Most teams that run this pass discover their portfolio is lopsided toward quadrant two: plenty of long guides, almost no lead answers. That's a two-week fix, not a content rebuild.
Where the two surfaces sit next to everything else
Both are Google surfaces, and both are only part of the picture. Our engine-comparison work found that cross-vendor agreement on brand picks runs lower than the 18.6% two-Google-surfaces number — how much ChatGPT, Perplexity and Gemini overlap on brand picks puts numbers on that spread. AI Mode is also the closest of Google's surfaces to the multi-step agent pattern that deep research modes run at longer horizons: more sub-searches, more sources, more weight on pages that resolve a whole sub-question in one passage.
The planning implication is narrow but real: treat AI Overviews as your Google-index proxy and AI Mode as your agent-retrieval proxy. Wins in the second category tend to travel to other agentic surfaces. Wins in the first mostly don't.
How to measure this without rankings or Search Console splits
Search Console folds AI feature traffic into overall Search totals — Google states that clicks and impressions from AI features are included in overall Search performance data, with no dimension separating AI Mode from AI Overviews from classic blue links. There's no rank to track and no filter to apply.
So measurement has to be observational. The workable method:
- Fix a prompt library. 80–200 real buyer questions, versioned. Changing prompts mid-quarter destroys comparability.
- Run both surfaces on the same schedule, same locale, same device class, logged out.
- Record URL-level citations, not just brand mentions — store both, because the 57.2% mention overlap and 18.6% link overlap answer different questions.
- Compute overlap per query type, never in aggregate. The aggregate number hides the commercial-query collapse.
- Re-run before and after every content change so you can attribute a citation to an edit rather than to platform drift.
Step five is where most programs break. Platform-side variance is large enough to swamp a real improvement if you only measure once. Two guards that work: hold everything except the edited page constant during a test window, and require a change to survive at least seven consecutive daily runs before you call it a win. Our own AI visibility tool runs exactly this loop daily, but the method matters more than the tooling — a disciplined spreadsheet beats an undisciplined dashboard.
The same discipline applies past Google. Copilot and Grok have their own retrieval stacks and their own trigger behaviour, and are routinely missing from B2B dashboards entirely — the AI engines B2B brands forget to track covers what to add and in what order.
How stable are these wins? A 21-day churn test
We re-ran a 120-query subset every day for 21 consecutive days to separate real citations from noise.
| Cohort | Median day-over-day URL churn |
|---|---|
| AI Mode source set | 44% |
| AI Overviews source set | 27% |
| URLs that were double winners on day 1 | 12% |
Nearly half of AI Mode's cited URLs changed from one day to the next. A single-day snapshot of AI Mode is close to meaningless — the uncomfortable truth behind a lot of published AI citation research, including studies with sample sizes far larger than ours.
The payoff: double-winner URLs churned at 12%, roughly a quarter of AI Mode's baseline volatility. Winning both surfaces doesn't just double exposure — it buys durability. Pages the two systems independently agree on are the pages that stay put.
That stability is what makes cross-surface agreement a better health metric than raw citation count. Report "URLs cited on both surfaces for 7+ consecutive days" as your headline number, with raw citation count as a secondary. It moves slower, and it stops your dashboard from swinging on retrieval noise.
What to do first if you only optimize for one surface
Ranked by cost-to-impact from what we observed:
- Fix lead answers on existing long guides. Cheapest, fastest, and moves the surface with the stronger carry-over (Overviews → AI Mode at 49.2%).
- Build or upgrade definition pages for your category's core terms. 34.1% overlap and an 88% Overviews trigger rate make these the highest-probability double wins per hour spent.
- Add comparison tables to head-to-head and shortlist pages. 57% of double winners have one; 24% of AI Mode-only pages do.
- Earn third-party corroboration on the two or three commercial queries that matter most. Slowest and least controllable, but it's the only lever that moves the 9.3%-overlap shortlist queries.
- Update your own pages last for Overviews purposes, first for AI Mode — owned content carries 21.3% of AI Mode citations and 8.7% of Overviews citations. If your goal is a specific tab, the ordering flips.
A 30-day plan to win both surfaces
- Days 1–5. Fix a 100-prompt library from real buyer questions. Baseline both surfaces daily. Edit nothing yet.
- Days 6–8. Tag every cited URL into the four quadrants. Segment by query type. Identify your quadrant-two pages — long guides with buried answers.
- Days 9–16. Rewrite the top 10 quadrant-two openings: a 40–60 word direct answer in the first 150 words, then one comparison table. Change nothing else, so attribution stays clean.
- Days 17–23. Extend the top 5 quadrant-three pages with 5–8 new self-contained sub-answers, each under a descriptive question-form H2.
- Days 24–30. Re-measure. Compare overlap per query type against baseline. Expect movement on definitional and how-to prompts first; shortlist queries typically lag a further 3–6 weeks because they depend on off-site corroboration you can't ship on a deadline.
Run the same loop next quarter on the next ten pages. Answer engine optimization compounds through repetition, not through one heroic rewrite.
Frequently asked questions
Is AI Mode replacing AI Overviews?
No evidence of that in our data. Over six weeks the Overviews trigger rate held within a 4-point band with no downward trend. They serve different jobs: Overviews compresses an answer onto the results page, AI Mode hosts a conversation in its own tab. Plan for both to persist.
If I can only optimize for one, which should I pick?
AI Overviews, on the evidence above. Half of Overviews citations also land in AI Mode, versus 23% in the other direction. Optimizing for the tighter filter tends to carry you into the looser one — the reverse rarely holds. The exception: if your category's commercial queries trigger Overviews under ~45% of the time, that carry-over has little to carry, and AI Mode is the better first target.
Does winning AI Overviews mean I'll also be cited by ChatGPT or Perplexity?
Not reliably. Google's two surfaces share 18.6% of URLs with each other; cross-vendor agreement is a different and generally lower number. ChatGPT and Perplexity brand mentions depend on separate retrieval and grounding stacks, which is why single-engine reporting misstates real coverage.
Why does AI Mode cite my own site more than AI Overviews does?
AI Mode's fan-out generates brand-specific sub-searches — pricing, integrations, support — where your own pages are the most direct source. Overviews weights independent corroboration more heavily, sending only 8.7% of citations to the queried brand's domain in our set.
How many prompts do I need before the numbers mean anything?
Given 44% daily churn in AI Mode, treat any conclusion from fewer than 50 prompts observed over fewer than 14 days as directional only. Overlap and share-of-voice figures stabilized in our data at roughly 80 prompts tracked across three weeks.
Can I show up in AI Mode without ranking in classic search?
It happens, but it's the exception. AI Mode's sub-searches still retrieve from Google's index, so pages that rank nowhere for any sub-question rarely surface. The realistic version is ranking modestly (positions 10–30) for several sub-questions rather than top-3 for the head term — breadth of adequate rankings beats a single strong one.
Do AI Mode citations send meaningful traffic?
Less per citation than a classic top-3 ranking, and the click happens further down the funnel. Treat AI Mode citations as an assisted-influence metric, not a traffic channel — which is why the mention column (57.2% overlap) belongs in reporting alongside the link column (18.6%).