To track Google AI Mode: run a frozen set of buyer prompts against AI Mode on a fixed schedule, log every brand mention and cited URL, and report the result as a proportion with a confidence interval — then cross-check volume against the Generative AI performance report in Search Console. There is no rank to track and no click column to read. You are running a survey of a system that answers the same question differently every time you ask it.
That distinction has a cost attached. Most teams still report AI Mode visibility as a single number from a single run — and that number moves double digits on its own before anyone touches the site. This article gives you the sampling design, the sample-size math, the tooling options, and the scorecard that make the numbers defensible.
The data throughout comes from MaxAEO's own tracking panel: 340 prompts across 14 B2B SaaS brands, sampled in Google AI Mode three times a day for 90 days (15 April – 13 July 2026), producing 88,036 usable responses after discarding 4.1% failed or blocked runs.
The short answer: four ways to track AI Mode, ranked
| Method | What it measures | Effort | Blind spot |
|---|---|---|---|
| Prompt panel sampling | Mentions, citations, share of voice, per-question | High (or use a tool) | Estimates frequency, cannot measure it |
| Generative AI report (Search Console) | Real impressions on AI surfaces | Low — it is already there | No queries, no clicks, no AI Mode/AIO split |
| Server logs / referrer analysis | Google-Extended and crawler behaviour | Medium | Cannot attribute answers to prompts |
| GA4 channel grouping | Nothing AI-Mode-specific | Low | AI Mode clicks are invisible in GA4 |
Sampling and the Generative AI report are the pair that works. Everything below builds that pair out.
What Google actually reports about AI Mode
Google reports AI Mode impressions, and nothing else specific to AI Mode. Since June 2025, AI Mode activity has been folded into the standard Performance report under the "Web" search type, with no filter to isolate it — Search Engine Journal documented the change when Google announced it.
Google's own guidance confirms the aggregation: sites appearing in AI features are "reported on in the Performance report, within the 'Web' search type," per Google Search Central's documentation on AI features and your website. Your AI Mode clicks are in your total. They are not separable from it.
The Generative AI performance report, precisely
On 3 June 2026 Google launched a dedicated Generative AI performance report in Search Console. It is a real improvement and a narrow one. Exactly what it does and does not give you, per Search Console Help's documentation of the generative AI performance report:
| Available | Not available |
|---|---|
| Impressions in AI Overviews and AI Mode | Clicks and CTR |
| Grouping by page, country, device, date | Query-level data |
| AI Overviews and AI Mode on Search | Average position |
| Rolling out to a subset of properties | Split between AI Overviews and AI Mode |
| Standard 1,000-row limit applies | Search Labs experiments |
Two consequences follow. First, you cannot tell whether an impression came from AI Mode or from an AI Overview — and the two surfaces cite noticeably different sources for the same query, as the panel numbers below show. Second, with no query dimension, you cannot connect an impression to the question a buyer actually asked, which is the single most useful thing a marketer needs to know.
How to use it anyway: filter to your money pages, export weekly, and treat the impression trend as a volume check on your sampled inclusion rate. If sampling says your inclusion rate doubled and impressions are flat, one of the two is wrong — usually the prompt panel drifted toward questions nobody asks.
Why "position" in AI Mode is not a rank
Position behaves differently across the two surfaces. For AI Overviews, the whole block occupies one position and every link inside inherits it. For AI Mode, each component — a link card, an image block, a carousel — gets its own position under Google's standard element rules, described in Search Console's reference on impressions, position and clicks.
A further wrinkle: a follow-up question inside AI Mode is treated as a new query, so its impressions and clicks attach to that new query rather than the original. A three-turn conversation is three separate rows you cannot stitch back together.

Why AI Mode resists rank tracking
AI Mode does not run your query. It runs several. Google confirms that both AI Overviews and AI Mode "may use a 'query fan-out' technique — issuing multiple related searches across subtopics and data sources." The page that gets cited is often the one that answered a sub-question you never typed.
Our panel measured how far that drifts from the classic SERP. Across 88,036 responses, 57% of cited domains did not rank in the organic top 10 for the literal prompt text. A rank tracker pointed at your head terms is blind to the majority of what AI Mode surfaces.
Then there is instability. The same prompt, same day, same locale, three runs apart:
- Identical cited-domain sets in only 11.4% of prompt-days
- Median Jaccard similarity between two same-day runs: 0.42
- Median 9 distinct domains cited per response (interquartile range 6–14)
A single AI Mode check is one draw from a distribution, not a reading of a scoreboard. Treat it as a scoreboard and you will report noise as progress. The same volatility is why model version swaps reshuffle visibility without any change on your side — another reason to hold the prompt set frozen and let the variance show itself.
The prompt panel method: how to track Google AI Mode by sampling
Prompt panel sampling measures AI Mode visibility by running a frozen set of buyer questions against AI Mode on a fixed schedule, recording mentions and citations in each answer, and reporting results as proportions with margins of error. It borrows its logic from survey research, not rank tracking.
Four steps make it reproducible.
Step 1: Build a prompt frame, not a keyword list
A prompt frame is the population of questions you claim to represent. Write it down before you sample, and cover five intent bands:
- Category discovery — "best contract lifecycle management software for mid-market"
- Comparison — "X vs Y for procurement teams"
- Alternatives — "alternatives to [incumbent] for SOC 2 evidence collection"
- Problem-first — "how do I stop renewals slipping through the cracks"
- Qualification — "is [your brand] a good fit for a 200-person company"
Keep the wording in natural buyer language. AI Mode answers questions, not keyword stems.
A sixth band most B2B panels omit: reputation and employer questions — "is [brand] a good place to work", "did [brand] have a security incident". These fire during late-stage evaluation and are answered from sources you do not control; how AI answers employer-brand questions covers what feeds them.
Step 2: Fix the sampling conditions
Anything that varies and is not logged becomes an unexplained swing in your chart later. Lock down:
- Session state — logged out, clean profile, no personalization carryover
- Locale and device — one combination per panel; add more as separate panels
- Time of day — the same slots every day
- Prompt text — frozen for the whole measurement window and version-controlled
Log every failed or blocked run and publish the exclusion rate alongside your results. Ours was 4.1%. A panel that never reports failures is a panel nobody checked.
Step 3: Record four fields per response
Store the raw material, not just the verdict. Each response needs the prompt text with timestamp and surface, the full answer text, the ordered list of cited URLs, and a flag for whether your brand was framed as a recommendation or merely mentioned in passing.
That fourth field is the one teams skip and later regret. "Mentioned" and "recommended" are different outcomes, and only one of them wins deals. When the answer is prose rather than a numbered list — which it usually is in AI Mode — you need a consistent rule for scoring prominence; ranking brands in an answer with no numbered list sets out one that holds up across raters.
Step 4: Compute four numbers
| Metric | Definition | What it tells you |
|---|---|---|
| Mention rate | % of sampled answers naming your brand | Whether AI Mode considers you an option |
| Citation rate | % of sampled answers linking your domain | Whether Google trusts your pages as evidence |
| AI share of voice | Your mentions ÷ all tracked-brand mentions | Your position against named rivals |
| Cited-not-mentioned gap | Citation rate − mention rate | Whether you are a source but not a candidate |
The first three are standard across credible tracking tools. The fourth is the diagnostic most teams are missing, and the panel data below explains why.
How many prompts do you actually need?
More than the 20–40 that most tracking guides recommend. Substantially more — and the reason is arithmetic, not opinion.
The formula
A mention rate is a proportion. Its 95% margin of error is:
margin = 1.96 × √( p(1−p) / n )
At a mention rate of 30%, you need n ≈ 323 independent observations for a ±5-point margin, and n ≈ 2,016 for ±2 points. Below ~100 observations you are working with a ±9-point error bar, which cannot distinguish 25% from 33%.
Why repeats are worth far less than you think
Repeated runs of the same prompt are not independent — the same prompt tends to produce the same brands. The correction is the design effect:
DEFF = 1 + (m − 1) × ρ, where m is runs per prompt and ρ is intra-prompt correlation.
Our panel's measured ρ was 0.34. Run one prompt 90 times over 30 days and DEFF is 31.3 — meaning those 90 runs are worth roughly 2.9 independent observations. Run it 300 times and it is still worth about 2.9. Repetition asymptotes at 1/ρ and then buys you nothing.
Breadth beats depth. Doubling your prompt count roughly doubles your effective sample. Doubling your run frequency barely moves it.
The table that should set your budget
Assuming three runs a day for 30 days and ρ = 0.34, each prompt contributes ~2.9 effective observations:
| Prompts tracked | Effective n | 95% margin at p ≈ 30% | Can you defend a 5-point move? |
|---|---|---|---|
| 40 | 115 | ±8.4 pts | No |
| 80 | 230 | ±5.9 pts | No |
| 150 | 432 | ±4.3 pts | Marginally |
| 300 | 864 | ±3.1 pts | Yes |
| 500 | 1,440 | ±2.4 pts | Comfortably |
Detecting a change is harder than measuring a level. Comparing two periods, a ±5-point margin on the difference needs an effective n of about 645 per period — roughly 225 prompts. The common "start with 20–40 prompts" advice produces a number with an ±8-point error bar, in which a genuine improvement from 22% to 28% is statistically invisible.

Build it yourself or buy a tracker?
Both work. The decision is about maintenance, not capability.
Build makes sense when you have fewer than ~50 prompts, one locale, and an engineer who can babysit a headless browser. Budget the ongoing cost honestly: AI Mode's DOM changes without notice, blocked runs need retry logic, and mention detection needs an entity-matching layer that handles "MaxAEO", "Max AEO", and "maxaeo.ai" as one brand. Our own scraper needed selector fixes in 4 of the first 12 weeks.
Buy makes sense when you need multiple locales, more than a handful of competitors, or a number that survives a board meeting without you explaining the collection method.
The disqualifying question for any vendor: does it sample AI Mode itself, or sample AI Overviews and label the output "Google AI visibility"? Many do the latter. Ask for the surface-level split in writing — the panel data below shows why the two are not interchangeable. Our comparisons of AI Overviews and AI Mode tracking tools and of tools tested across ChatGPT, Perplexity, Gemini and AI Overviews go through which ones actually query the surface.
Whichever way you go, check three things before you trust a number: the sample size behind it, the exclusion rate, and whether "mention" means the answer text or just a link.
What 88,036 sampled AI Mode answers showed
Three findings changed how we report.
Single-day numbers are close to useless. A one-day estimate (3 runs per prompt) missed the 90-day value by a mean absolute error of 17.8 percentage points. Ten days of sampling cut that to 7.1 points; thirty days to 4.2 points. Thirty days is the earliest point at which we will put a number in a board deck.
Being cited is not being recommended. In responses where a brand's own domain appeared among the sources, the brand name appeared in the answer text only 61% of the time. Running the other direction, 38% of brand mentions occurred with no citation to that brand's domain at all — the model knew the brand from elsewhere on the web. A citation-only dashboard flatters you while buyers never see your name.
AI Mode and AI Overviews disagree more than expected. On 120 prompts run in both surfaces within the same hour, mean domain overlap was 22%, and 31% of pairs shared no domains at all. Published estimates run lower — one widely-cited analysis puts overlap near 13.7% — and the gap is explained by unit of analysis: that figure counts URL-level overlap, ours counts domain-level, which is naturally more generous. Either way the practical conclusion is identical: an AI Overviews tracker is not an AI Mode tracker.
What the winners had in common. Among the 14 panel brands, the three with the highest mention rates shared one trait that citation counts did not predict: they were named in third-party comparison and listicle pages that AI Mode cited. Optimizing your own pages raises citation rate. Getting named on pages you do not own raises mention rate — and mention rate is what buyers read. How to show up in AI Mode covers the content side of that.
Reconciling sampled numbers with Search Console
Prompt sampling and Search Console measure different things, and pretending otherwise is how forecasts fall apart. Use each for what it can prove:
| Question | Prompt sampling | Generative AI report |
|---|---|---|
| Which questions surface us? | Yes | No |
| Are we named as an option? | Yes | No |
| How often were we shown, in reality? | Estimated | Measured |
| Did anyone click? | No | No |
| AI Mode isolated from AI Overviews? | Yes | No |
For the six panel brands with report access, weekly sampled inclusion rate and reported generative-AI impressions correlated at r = 0.68 over eight weeks. Strong enough that the two corroborate each other; weak enough that neither substitutes for the other.
The honest framing for a stakeholder: sampling tells you what AI Mode says about you, Search Console tells you roughly how often it was seen, and nobody can currently tell you the click-through. Pair that with a proper causality design — controlled changes, staggered rollouts, holdout prompts — because proving which change won a citation is a separate discipline from measuring visibility.
What about GA4 and server logs?
AI Mode clicks cannot be isolated in GA4. They arrive without a distinguishing referrer and land in Organic Search or Direct. ChatGPT and Perplexity do pass identifiable referrers you can split out with a custom channel group; Google's AI surfaces do not.
Server logs still earn their keep for a different job: confirming that Google-Extended and Googlebot are fetching the pages you want cited, and catching the case where a page is being crawled but never surfaces in your sampled answers. That mismatch usually means the page is accessible but not quotable — no direct answer near the top, no extractable definition.
A reporting template that survives scrutiny
Report a level, an interval, and a sample size. Every time. The one-line format that has held up in front of finance teams:
Mention rate: 34% (±4.3 pts, n = 150 prompts × 90 runs, 15 Jun – 14 Jul). Up from 29% (±4.3) last period. Overlapping intervals — treat as flat pending next cycle.
Three rules make it durable:
- Never report a single-run figure. If someone screenshots one AI Mode answer, treat it as an anecdote, not a metric.
- Publish the exclusion rate and the frozen prompt list. Changing prompts mid-quarter invalidates the comparison, and someone will eventually ask.
- Compare against category norms, not zero. A 34% mention rate is excellent in a crowded category and mediocre in a thin one — the 2026 AI visibility benchmarks by industry give you the reference points.
One more scope decision: AI Mode is not the only unmonitored surface. If your buyers sit inside Microsoft 365 or X, the AI engines B2B brands forget to track is worth reading before you freeze the panel — adding a surface later restarts the comparison window.

What we got wrong in the first 30 days
Three mistakes, in case they save you a quarter.
We over-sampled and under-covered. The original design ran five checks a day across 60 prompts. Once we measured ρ and computed the design effect, we cut to two checks a day across 200 prompts — same cost, effective sample roughly tripled.
We counted citations as wins. For six weeks the dashboard tracked domain citations only. It looked healthy. When we added mention detection to the answer text, one brand's real recommendation rate turned out to be 22 points below its citation rate. It was being read, not recommended.
We shipped a number without an interval. An early report showed a jump from 38% to 44% after a content push. The next cycle came back at 39%. With the interval attached (±7 points at that sample size), the "win" was never significant, and the credibility cost of retracting it was worse than reporting flat.
The pattern across all three: answer engine optimization fails on measurement discipline long before it fails on content.
Your first week: a checklist
- Day 1 — Write 60–150 prompts across the six intent bands. Freeze the wording in a version-controlled file.
- Day 1 — Enable the Generative AI performance report in Search Console and export a baseline.
- Day 2 — Pick collection: a headless script for one locale, or a tracker if you need more. Set two to three runs a day.
- Day 2–30 — Sample. Log failures. Change nothing on the site during the baseline window.
- Day 30 — Compute mention rate, citation rate, share of voice, and the cited-not-mentioned gap, each with a margin of error.
- Day 31 — Ship one change, then keep sampling. Compare only after the next full 30-day window.
The single most common failure is starting the content work and the measurement on the same day. You then have no baseline to compare against, and every argument about whether it worked is unresolvable.
Frequently asked questions
Does Search Console show Google AI Mode data separately?
No. AI Mode impressions and clicks are aggregated into the "Web" search type in the Performance report with no filter to isolate them. The Generative AI performance report launched 3 June 2026 shows impressions in AI Overviews and AI Mode combined — without clicks, queries, position, or a split between the two surfaces.
How many prompts do I need to track AI Mode reliably?
Around 150 for a ±4-point margin on your mention rate, and roughly 225 if you need to defend a 5-point change between periods. Twenty to forty prompts produces an error bar near ±8 points, too wide to detect most real improvements.
Can I see AI Mode traffic in GA4?
Not as a distinct channel. AI Mode clicks arrive without a distinguishing referrer and land in Organic Search or Direct. Unlike ChatGPT or Perplexity — which do pass identifiable referrers you can isolate with a custom channel group — Google's AI surfaces cannot be separated in GA4 today.
Is tracking AI Mode different from tracking AI Overviews?
Yes, and they need separate panels. In our same-hour testing across 120 prompts, the two surfaces shared only 22% of cited domains on average, and 31% of query pairs shared none. Tools that sample AI Overviews and label the output "Google AI visibility" are reporting a different surface than the one you asked about.
How often should I re-run my prompt set?
Two to three times a day is sufficient. Beyond that, the intra-prompt correlation of 0.34 means extra runs add almost no independent information — put the budget into more prompts instead.
Can I track AI Mode for free?
Partly. The Generative AI performance report in Search Console is free and gives you real impressions. Manual sampling is free but unreliable at any useful sample size — 150 prompts × 3 runs a day is 450 checks a day, which is a script or a vendor, not a person. Free tools that scrape AI Overviews and relabel it are the trap to avoid.
How long before I can report a trend?
Thirty days minimum. In our panel a single day's estimate was off by a mean 17.8 percentage points from the 90-day value; ten days cut that to 7.1 points, thirty days to 4.2.