The first AI visibility report we bought looked great.
There was a score, a competitor chart, a list of prompts, and enough week-over-week movement to fill a monthly deck. Three months later, our content meeting still ended with the same question:
What are we actually doing on Monday?
I work on growth/content for a SaaS company. We had no shortage of data. We knew where competitors appeared, which engines mentioned us, and which prompts looked weak. The problem was that the dashboard stopped exactly where the work started.
We kept running into three versions of the same failure.
First, the overall visibility score moved, but nobody could explain why. One engine could be improving while another was dropping, and the average hid both changes.
Second, we got a content-gap list with advice like “publish more authoritative content.” Fine, but where? Our own blog? A comparison page? An industry site already being cited for that prompt? A community discussion? Those are completely different jobs.
Third, after we shipped something, we couldn’t prove whether it helped. A mention rate might move a month later, but the model had changed, competitors had published, and the prompt set was different. We were doing attribution by vibes.
There was also an earlier problem we missed: sometimes the model had us in the wrong category. If an AI system thinks you are an analytics add-on when you are actually workflow software, publishing more pages around the wrong framing can make the confusion worse.
So we stopped comparing tools by the number of charts and rewrote the buying decision as three questions.
1. Can it tell me where the missing answer currently comes from?
Not “your content needs more authority.” I mean a specific source pattern.
For a weak prompt, I want to see the pages the engines actually cite, grouped by source and engine. If the answer is consistently coming from review sites, Reddit threads, or competitor comparison pages, that changes the next action.
The test is simple: open one missing prompt and ask the tool to show the cited sources, the exact passages being pulled, and a proposed placement. If the recommendation is still just “write a blog post,” it hasn’t crossed the gap from monitoring to execution.
2. Can it separate mention, recommendation, and category understanding?
A brand mention is not the same as a recommendation. And a recommendation is not useful if the model describes the company incorrectly.
We now check the same prompt across engines and look for three different things:
- Did the brand appear at all?
- Was it recommended for the use case, or merely named?
- How did the answer categorize the product and its audience?
This is where a single blended score becomes actively unhelpful. ChatGPT, Gemini, Perplexity, and Google’s AI surfaces can disagree. I would rather have eight imperfect engine-level views than one clean average that hides the disagreement.
The on-screen test: pick a prompt where you know your positioning. If the tool cannot show the actual answers and let you compare how each engine describes the company, the score is not diagnostic enough to guide content.
3. Can it preserve a baseline and verify the same work later?
Before anything ships, we save the prompt set, engine set, cited URLs, mention state, recommendation state, and category description. Then we log the action: what was published, where, and when.
Later, we rerun the same prompt set against the same engines. Not a new dashboard view with a new denominator. The same battlefield.
The test here is boring but important: can you click from an action to its before state and its later result? If the system cannot connect diagnosis, publication, and verification, the team will eventually reconstruct the story in a spreadsheet before every budget review.
None of this proves causality perfectly. AI answers change, competitors move, and a citation can appear for reasons you did not control. I don’t think any dashboard can honestly remove that uncertainty.
But it does give us a much better standard than “the score went up.” We can say which gap we targeted, why we chose a particular source, what changed in the same prompt set, and what still did not move.
For teams doing this seriously: after you ship a batch of AEO work, what evidence do you use to decide whether it worked well enough to repeat?
