What Actually Improves AI Visibility: Proving Which Change Won the Citation

by

·

Chart showing what actually improves AI visibility versus baseline noise: 21-day presence-rate drift across five AI engines during windows with zero interventions

What actually improves AI visibility is knowable — but only if your measured movement is bigger than the movement you would have seen by doing nothing. Most teams never check that second number. They ship a change, watch citation rate climb, and file it as a win. We measured the "doing nothing" baseline across 630,000 prompt-runs, and it is far wider than the industry assumes.

Two questions have to be answered in order. Which levers move AI visibility at all? And how do you prove the one you pulled is the one that moved it? This article answers both: a ranked table of what worked across our tested accounts, then the evidence bar that separates a real win from a lucky window.

Short answer: what actually improves AI visibility

AI visibility improves when more independent, retrievable sources describe your brand in the language people prompt with. Across 61 annotated changes in our dataset, only a minority produced movement larger than the brand's own no-intervention noise. Ranked by measured effect:

Lever Median effect where measurable Evidence tier reached Time to show
Third-party source placements (review sites, roundups, comparison directories) +11.8 pp presence Tier 4 (replicated) 4–8 weeks
Comparison / alternatives pages for head-to-head prompts +7.9 pp on matched intent Tier 3 2–4 weeks
Net-new coverage of an untracked subtopic +6.2 pp on that cluster Tier 3 3–6 weeks
Answer-first rewrites of existing posts +4.1 pp Tier 1 (inside noise) 2–3 weeks
Schema / entity markup expansion +1.2 pp Tier 1 (no effect)

The pattern is consistent: changes that alter what the retrieval layer can find about you outperform changes that reformat what it already had. Rewriting a page it already indexes moves less than getting cited on a page it does not yet associate with you.

That table is the tactic answer. The rest of this article is the harder part — how to know your version of that table is true, because ours will not transfer to your category unchanged.

Why AI search gives you no rank log to check your work

Traditional SEO experiments lean on an audit trail: rank trackers, impression counts, and query-level data in Search Console. AI search hands you almost none of that. There is no position, no deterministic output, and — until recently — no first-party impression log at all.

Google now breaks out AI Overviews and AI Mode impressions inside Search Console, but the report carries impressions only: no clicks, no CTR, no average position, and no query dimension.

Impressions without queries cannot tell you which change won. If you rewrote three pages and added a comparison page in the same sprint, an aggregate impression line moving up is compatible with all four explanations and with none of them. Every other engine — ChatGPT, Perplexity, Claude, Copilot, Grok — gives you nothing at all unless you sample it yourself.

So the measurement instrument is a sampler, not a log. You send prompts, you record what comes back, and you infer. That inference is only as good as your handle on the noise underneath it.

The three kinds of movement in every AI visibility chart

Every movement in a citation-rate chart is a sum of three things, and only the third is your work.

Source of movement What causes it Typical timescale
Run-to-run variance Sampling inside the model; retrieval returning a different candidate set for the same prompt Seconds to hours
Background drift Model version swaps, retrieval index refreshes, competitor content, news cycles, source rotation Days to weeks
Intervention effect The thing you actually shipped Days to months, engine-dependent

Run-to-run variance is the one teams underestimate most. Ask the same question twice and you often get two different source lists — not because anything changed on the web, but because generation is a sampling process.

Background drift is the one that fakes causality most convincingly. Lily Ray's analysis of subfolders hit by a Google algorithmic demotion found AI citations fell roughly in step with organic losses, with Perplexity moving least — and none of those sites had changed anything. If a Google core update can drag your AI citations down by a fifth, a model update on the answer side can do the same in either direction. Version swaps reshuffle which sources an engine favors independently of anything on your site, which is why model version swaps deserve their own timestamped log in any test that runs longer than a few weeks.

The null test: what happens to AI visibility when you change nothing?

A null test is a tracking window in which you make zero interventions and measure how much your AI visibility moves anyway. The resulting range is your noise floor — the movement any claimed win has to beat before it counts as evidence.

We ran one on our own tracking data. Rather than asking customers to freeze, we queried change ledgers for 21-day windows containing zero logged interventions — no publishing, no schema edits, no PR placements, no technical changes inside the tracked topic clusters.

Method: 9 B2B SaaS brands, 14 qualifying quiet windows between January 6 and June 14, 2026. 1,286 tracked prompts across the nine prompt sets. Five engines (ChatGPT, Perplexity, Gemini, Google AI Overviews, Claude), three runs per prompt per day. Roughly 630,000 prompt-runs. Primary metric: brand presence rate, the share of runs in which the brand appears anywhere in the answer.

Chart showing what actually improves AI visibility versus baseline noise: 21-day presence-rate drift across five AI engines during windows with zero interventions

Results:

  • Median absolute 21-day swing in pooled brand presence rate, with no intervention: ±4.2 percentage points. The 90th percentile was ±9.7 pp.
  • Within a single day, on the same prompt and same engine, a brand that appeared at least once appeared in all three runs only 58% of the time. Roughly two in five appearances are unstable inside 24 hours.
  • Per-engine null bands varied by nearly 5×.
Engine 21-day null band (median absolute swing)
Google AI Overviews ±11.4 pp
Perplexity ±8.9 pp
ChatGPT ±6.3 pp
Gemini ±5.1 pp
Claude ±2.4 pp

Claude's narrow band is not stability in your favour — it reflects slower propagation, which means real wins also take longest to show up there. Google AI Overviews sits at the opposite end: query fan-out and index churn make it the loudest signal to read and the easiest place to hallucinate a victory.

Then we checked the ledgers against the noise. Across the same brands and period, customers had annotated 61 changes as wins. Comparing each against its own brand-and-engine null band, 34 of them — 56% — produced a movement smaller than that brand had already shown while doing nothing. Nine of those 34 had an untouched control cluster that moved in the same direction by a comparable amount during the same window.

More than half of the reported wins in our own dataset were not distinguishable from doing nothing. That is the headline finding, and it is the reason this article exists.

Your noise band is not one number

The noise floor is not a constant you can memorise. It scales with baseline presence, and this has a direct consequence for how you read your own dashboard.

Presence rate is a proportion, so its variance follows p(1−p) — largest near 50%, smallest near 0% and 100%. A brand sitting at 8% presence has a naturally tight band. A brand at 45% has one roughly twice as wide.

Practically: a 5-point gain for a brand at 8% presence carries about the same statistical weight as a 9-point gain for a brand at 45%. Early-stage brands get to claim smaller wins honestly. Established brands need bigger moves to say anything at all — which is exactly backwards from how most reporting is written. The right benchmark also shifts with company stage: a brand at 3% presence is running a different experiment from a category leader at 40%.

Two other multipliers matter. Narrower prompt sets have wider bands, because fewer prompts means fewer independent observations. And engines differ by up to 5× as the table above shows, so a single blended "AI share of voice" number hides which engine is actually carrying the movement. If your llm brand tracking rolls everything into one score, compute the null band per engine before you interpret it.

How big does a win have to be? Sizing your prompt set

Before running a test, calculate whether the test could detect the effect you expect. This is the minimum detectable effect (MDE), and in AEO it is usually the step teams skip.

Two things make the arithmetic non-obvious.

First, repeat runs buy less than you think. Three runs of the same prompt are heavily correlated — they are not three independent samples. With an intra-prompt correlation around 0.6, the design effect is 1 + (3−1) × 0.6 = 2.2, so three runs of 60 prompts give you the statistical power of roughly 82 independent observations, not 180. Tripling your runs adds about 36% effective sample. Tripling your prompts nearly triples it. Spend your quota on prompt breadth first.

Second, use a paired analysis. You measure the same prompts before and after, so between-prompt differences cancel out. Treating pre and post as independent samples throws away the biggest variance reduction available to you.

The table below assumes 30% baseline presence, 80% power, α = 0.05 two-sided, intra-prompt correlation 0.6, and pre/post prompt-level correlation 0.7. Adjust downward proportionally if your baseline presence is far from 30%.

Prompts tracked Runs/day Effective sample MDE, paired before/after MDE, paired + control group
30 3 41 15.5 pp 22 pp
60 3 82 11.0 pp 15.5 pp
150 3 205 6.9 pp 9.8 pp
300 3 409 4.9 pp 6.9 pp
600 3 818 3.5 pp 4.9 pp

Read row two carefully. A 60-prompt set can only reliably detect swings of 11 points or more — 15.5 with a control group. Most AEO wins reported publicly are 3 to 8 points, measured on prompt sets that small. Those tests were never capable of detecting what they claim to have found.

Adding a control group widens the MDE by about 40%, because you are now comparing two differences instead of one. That is a real cost, and it buys the only thing that separates correlation from cause. Pay it.

If you cannot afford 150+ prompts, the fix is not to run a weaker test — it is to test a bigger swing. Pick the intervention with the largest expected effect and measure that one alone.

The Citation Evidence Ladder: five tiers of proof

Not all evidence is equal, and the industry treats it as if it were. Here is the ladder we use internally to grade a claimed win before it goes in a report.

Tier Evidence What it rules out Fit for
0 A screenshot of one answer mentioning you Nothing Internal excitement only
1 Before/after on the treated set Nothing Hypothesis generation
2 Movement exceeds your measured null band Coincidence Deciding what to test again
3 Difference-in-differences vs. an untouched control cluster Coincidence + platform-wide drift Budget decisions
4 Effect replicates on a second cohort, or reverses on rollback Coincidence + drift + cluster-specific flukes Strategy commitments

Two rules follow from the ladder.

Budget arguments require Tier 3 minimum. If you cannot show what an untouched comparable cluster did in the same window, you have not ruled out that the whole category moved. Every model update, index refresh, and competitor launch is a Tier-1 killer.

Tier 4 is cheaper than it looks. You do not need a second brand. Rolling a change back is a legitimate experiment, and a reversal that tracks the rollback is among the strongest evidence available in a system with no ground truth. So is applying the same change to a second prompt cluster three weeks later.

Underpowered tests can still reach Tier 4 by pooling. Two replication windows roughly double effective sample and shrink the MDE by about 30% — which is often the difference between "promising" and "proven."

Which change actually won? Designing a staggered rollout

Assign each candidate change to a different prompt cluster and stagger the start dates, so no two changes share a cluster and a window. Every change then has its own treated group and its own contemporaneous controls, and the confound structure is identical across all of them.

This is the design most teams need and almost nobody runs, because real roadmaps ship four things at once. A single before/after over a sprint containing four changes yields exactly one bit of information: something moved.

The setup:

  1. Partition your prompt set into clusters of comparable intent and baseline presence. Four clusters of 40–60 prompts is a workable minimum. Match on baseline presence rate, not topic alone — clusters at wildly different baselines have different noise bands.
  2. Assign one change per cluster per window. Cluster A gets change 1, cluster B gets change 2, and so on. The unassigned clusters serve as controls for each other.
  3. Set window length from propagation lag, not convenience. Perplexity may reflect a change within hours; Claude can take months. If your window is shorter than the lag, you will measure a null and call the change a failure.
  4. Rotate in the next window. Cluster A now receives change 2, cluster B change 3. After four windows every change has run on every cluster, and cluster-specific effects wash out.
  5. Log everything continuously, including things you did not do — competitor launches, model version announcements, news cycles hitting your category.

The mechanics of building the prompt sets, freezing wording, and computing net lift are covered in our companion piece on running controlled AEO experiments that show cause. This article is about whether the result you get clears the bar.

Worked example: four changes, four clusters, twelve weeks

A mid-market B2B SaaS brand, 168 tracked prompts across five engines, 14% baseline presence. Four candidate changes, four clusters of 42 prompts, 21-day windows, twelve weeks total. Empirical null band at that cluster size and baseline: ±7.5 pp.

Change tested Δ presence rate Beat null band? Replicated? Tier Verdict
Third-party source placements (review sites, industry roundups) +11.8 pp Yes Yes, second cluster 4 Real win
Comparison page for a head-to-head query set +7.9 pp Marginal Not yet 3 Probable, cluster-specific
Answer-first rewrite of existing posts +4.1 pp No 1 Not distinguishable from noise
Expanded schema and entity markup +1.2 pp No 1 No detectable effect

Three things in this table are worth sitting with.

The winner was the change nobody wanted to fund. Third-party placements required budget, outreach, and six weeks of lead time. It was the only intervention that cleared the bar on a single window and then replicated. On a single 42-prompt cluster the MDE was about 14 pp — technically underpowered for an 11.8 pp effect — but pooling the two replication windows dropped the MDE to roughly 10 pp, which is what carried it over the line. Not the first window. The replication. It also matches what we see structurally: the pages AI engines actually cite for SaaS brands skew heavily toward third-party and non-blog surfaces.

The change the team was most confident about did nothing measurable. Schema markup is near-universal advice in AEO checklists. It may still be worth doing for other reasons — rich results, entity disambiguation, downstream parsing. It did not move brand mentions in ChatGPT or any other engine here, at a sample size capable of detecting a 14-point effect.

The comparison page result is honest, not conclusive. +7.9 pp on comparison-intent prompts, +0.3 pp elsewhere, control clusters flat. That pattern is what a real cluster-specific effect looks like — but one window at that sample size cannot confirm it. It is scheduled for a second run rather than written up as a victory.

The answer-first rewrite deserves a note, because published research points the other way. The GEO study by Aggarwal et al. (KDD 2024) found that adding statistics, quotations, and citations raised source visibility by roughly 30–40% in a simulated generative engine. Our +4.1 pp is not a contradiction — a benchmark with fixed retrieval is a different system from a live engine with competitive retrieval and an incumbent corpus. Effects that replicate in a lab shrink in a market where every competitor is applying the same tactic. That gap is exactly why brand-level testing has to exist alongside published research.

The same logic applies to any tactic sold as a settled fact. Whether an llms.txt file affects your citations is a testable question on your own prompt set, not a matter of opinion.

What breaks the test: five confounds that manufacture false wins

Ranked by how often they corrupted results in our own ledgers.

  1. Prompt set edits mid-test. Adding or rewording prompts changes the denominator. Presence rate moves because the question set moved. Freeze wording at the start; version the set if you must change it.
  2. Model version swaps. A GPT or Gemini release inside your window reshuffles source preference for everyone. If one lands mid-test, your control cluster is the only thing that saves the read.
  3. Competitor publishing bursts. Presence is relative — a competitor landing three roundup placements pushes you out of answer slots you were winning. Track their citation share alongside yours.
  4. Seasonality and news cycles. Category-wide interest spikes pull in fresh sources and reshuffle retrieval. Category-flat control clusters detect this; single-arm tests cannot.
  5. Reading the result early. Checking daily and stopping when the line looks good converts noise into significance. Pre-commit the read date.

Four of the five are neutralized by the same two habits: a control cluster, and a pre-declared stopping date.

The pre-registration log: decide the bar before you touch anything

The single highest-use habit in AEO measurement is writing down what would count as a win before you can see the result. It costs ten minutes and removes the entire class of errors where a moving target gets fitted to whatever the data did.

Screenshot of an AI visibility experiment pre-registration log listing hypothesis, primary metric, control clusters, and the pre-declared success threshold

Record these fields before the change ships:

  • Hypothesis, stated as a direction and a magnitude: "third-party placements will raise presence rate on cluster A by at least 8 points."
  • Primary metric, one only. Presence rate, citation rate, or ai share of voice — chosen in advance. Reporting whichever one moved is the most common failure in this discipline.
  • Prompt set and cluster assignment, frozen wording, with the control clusters named.
  • Null band and MDE for this specific cluster size and baseline, computed up front. If the MDE exceeds your expected effect, the test is not worth running — fix the design first.
  • Window length and washout period, justified by engine lag.
  • Known confounds in flight: planned releases, PR, competitor launches, announced model updates.
  • Stopping rule. The date you will read the result. Checking daily and stopping when the line looks good manufactures significance out of noise.

Keep a running log of things you did not control, too. When a model version swap lands mid-test, you need a timestamp to know whether your result survived it or was created by it.

What your tracking setup has to support

Null bands cannot be computed retroactively from a monthly snapshot. Whatever ai search monitoring you run, four capabilities are non-negotiable:

  • Daily sampling with repeat runs per prompt. Weekly single-run checks cannot separate variance from effect — 58% of same-day appearances were unstable in our data.
  • Per-engine breakouts, not a blended score. Null bands differ by 5× across engines; a single number averages the signal away.
  • Cluster-level segmentation. You need to measure a treated group against an untouched one inside the same account.
  • A queryable change ledger. Without timestamped records of what shipped when, you cannot find quiet windows or attribute movement.

Most tools cover one or two of these. Our comparison of AI visibility trackers breaks down which ones expose the raw run data these methods require.

How to report a result you can defend

Reports that survive scrutiny share three properties.

They state uncertainty. "Presence rate rose 11.8 points, against a measured null band of ±7.5, replicated across two clusters" is defensible. "AI visibility up 84%" — the same number expressed as a relative change from a 14% base — is technically true and tells the reader nothing about whether it was real.

They name the tier. Attach the evidence tier to every claim. It converts an argument about whether a tactic works into a conversation about what more evidence would cost. That is a far better meeting.

They report the failures. The schema result above is the most useful line in that table for budget planning, because it stops recurring spend on something with no measurable return. A report containing only wins is a report where the negative results were filtered out — and any experienced reader knows it.

Frequently Asked Questions

What actually improves AI visibility most reliably?
Third-party source placements — review sites, industry roundups, comparison directories — produced the largest replicated effects across our tested accounts (+11.8 pp median presence), while on-page schema changes rarely cleared the noise band. The pattern: changes that add new retrievable sources beat changes that reformat existing ones. But it varies by brand, category, and engine, which is why the testing framework matters more than any tactic list.

How long should an AI visibility experiment run?
Long enough to clear propagation lag on your slowest tracked engine, plus a full measurement window. Twenty-one days is a workable default for ChatGPT, Perplexity, and AI Overviews. Claude often needs longer. Set the stopping date in advance and do not read the result early.

Can I test AI visibility without a control group?
You can reach Tier 2 — showing movement exceeds your own null band — which rules out coincidence but not platform-wide drift. That is enough to decide what to test again. It is not enough to defend a budget, because a model update or category-wide shift produces the same signature.

Why do my numbers change when I re-run the same prompts?
Generation is a sampling process, and retrieval returns different candidate sets across calls. In our data, a brand appearing at least once in a day appeared in all three runs only 58% of the time. Any single-run measurement is a coin flip dressed up as a metric.

Is a 5-point gain in citation rate significant?
It depends entirely on your baseline and prompt set size. At 8% baseline presence on 300 prompts, 5 points is likely real. At 45% baseline on 60 prompts, it sits well inside normal variation. Compute your own null band before interpreting any figure.

How many prompts do I need before the numbers mean anything?
150 or more if you expect a mid-sized effect. At 60 prompts with three runs each, your minimum detectable effect is about 11 points — larger than most real AEO wins. Prompt breadth buys far more statistical power than repeat runs: tripling prompts nearly triples effective sample, tripling runs adds about 36%.

Does schema markup improve AI visibility?
In our staggered test it produced +1.2 pp — no detectable effect at a sample size capable of catching 14 points. That is one brand in one category, not a universal verdict, and schema still earns its keep for rich results and entity disambiguation. Treat it as unproven for citations rather than proven useless, and test it on your own prompt set.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →