Published methodology
The GEO method: how to audit whether AI answers cite you
Classic SEO asks whether you rank. Answer engines do not rank. They retrieve a handful of sources and synthesize one answer, so a site can sit at the top of page one and never be quoted. This methodology audits the second thing.
The four phases, the prompt set, the crawler matrix and the citability checklist below are published in full. They are free to read without an account, and you can run all of it by hand.
Last updated: 2026-09-07
Four phases, run in order
Do not skip Phase 1. Without a citation baseline, everything after it is speculation.
- Phase 1
Citation baseline
Run 10 to 15 prompts a real buyer would type. For each, record whether the brand is named at all, whether it is cited with a link or merely mentioned in prose, which domain the citation points to, and which competitors appear in what order. Report mention rate, cited-with-link rate and share of voice.
- Phase 2
Crawler access
Resolve robots.txt for each agent below, one at a time. One line in robots.txt outranks any amount of content work, so findings here take priority over everything else in the audit.
- Phase 3
Content citability
Take the three pages you most want cited and judge them against the checklist below. A retriever pulls a chunk, not a page, so the test is whether any 200 to 300 word block still makes sense with no surrounding context.
- Phase 4
Where citations come from
Tally which domains the answers actually cite: your own site, review sites, forum threads, roundup articles. This is what decides whether the next move is content on your own domain or presence on someone else's.
The buyer-intent prompt set (Phase 1)
Build 10 to 15 prompts a real buyer would type, covering all four intents. A set that is all category queries will overstate visibility.
| Intent | Shape | Example |
|---|---|---|
| Category | "best X for Y" | best expense tools for seed-stage startups |
| Comparison | "A vs B" | Ramp vs Brex for a 30-person team |
| Alternative | "alternatives to A" | alternatives to Expensify |
| Problem | symptom, no brand named | how do I stop chasing receipts from my team |
For each prompt, record four things: whether the brand is named at all; whether it is cited with a link or merely mentioned in prose; which domain the citation points to (your own site, or a third party such as a review site, forum thread or roundup); and which competitors appear, in what order. Report mention rate, cited-with-link rate and share of voice.
The AI crawler matrix (Phase 2)
The single most common cause of total absence from AI answers is an accidental block — and the agent most teams block is not the one that controls citations. Resolve each agent with standard robots.txt matching: the most specific User-agent group naming the agent wins, and * applies only when no group names it. Within the winning group the longest matching path rule wins, and Allow beats Disallow on an equal-length match.
| Agent | Operator | Purpose | What blocking it actually costs |
|---|---|---|---|
OAI-SearchBot | OpenAI | search index | citations in ChatGPT Search |
ChatGPT-User | OpenAI | live fetch during a chat | the model cannot open your page when a user asks about it |
GPTBot | OpenAI | training | background model knowledge, not search citations |
PerplexityBot | Perplexity | search index | Perplexity citations |
Perplexity-User | Perplexity | live fetch during a query | live page reads |
ClaudeBot | Anthropic | index and training | Anthropic-side retrieval |
Googlebot | main index | AI Overviews and AI Mode, plus normal search | |
Google-Extended | Gemini grounding and training | Gemini grounding only - not AI Overviews | |
Bingbot | Microsoft | Bing index | Microsoft Copilot, which rides the Bing index |
Applebot | Apple | index | Apple search surfaces |
Applebot-Extended | Apple | training | Apple Intelligence training only |
CCBot | Common Crawl | open crawl corpus | an input to many downstream models |
Report one row per agent with the verdict (ALLOWED / BLOCKED / PARTIAL), the rule responsible, and the impact. The rule responsible must quote the literal line from robots.txt, or say "no matching rule - allowed by default". Never state a verdict without the line that produced it.
The citability checklist (Phase 3)
Take the three pages you most want cited and judge each against the properties that actually get a passage lifted into an answer.
- Self-contained passages
- A retriever pulls a chunk, not a page. Can any 200 to 300 word block be quoted with no surrounding context and still make sense?
- A direct answer near the top
- Pages that open with positioning copy get skipped. The answer should appear in the first paragraph under the heading.
- Question-shaped headings
- Headings phrased as the question a user actually asks match retrieval far better than clever headings.
- Specifics
- Numbers, dates, named limits and prices are quotable. "Industry-leading performance" is not.
- First-hand evidence
- Original data, benchmarks and named methodology survive summarization. Restated common knowledge does not.
- Freshness signals
- A visible last-updated date, and content that is actually current.
- Structured data
- Organization, Product, FAQPage, Article. Verify it parses. Markup that renders is not necessarily markup that validates.
Limits of this methodology
- A single run is a snapshot, not a trend.
- Answer engines re-rank continuously and the same prompt can return different sources hours apart. To make the numbers mean anything, freeze the prompt set, re-run it on a fixed schedule, and record every result.
- Never state a citation rate you did not measure.
- An unfetchable page is a finding, not a gap to fill with a guess.
- Crawler names change.
- Before finalizing, check each operator's own published crawler documentation for agents added or renamed since this list was written, and state which list you used.
Frequently asked questions
- Do I need an account or a subscription to use this methodology?
- No. The prompt set, the crawler matrix and the citability checklist are published on this page in full and are free to read without an account. You can run the whole thing by hand.
- How is GEO / AEO different from classic SEO?
- Classic SEO asks whether you rank. Answer engines do not rank — they retrieve a handful of sources and synthesize one answer, so a site can sit at the top of page one and never be quoted. This methodology audits the second thing.
- Is this worth doing if we don't rank well yet?
- Yes, and that is where the opening is. Published research on AI citations consistently finds low overlap between the URLs cited in AI answers and the top organic results — most cited sources rank outside the top three. Being cited does not require ranking first.
- Should we just block AI crawlers to be safe?
- That depends on your position, and the decisions are separate: allowing search crawlers is what makes citation possible, allowing training crawlers affects model knowledge but not citation, and the two are independent. Publishers with a licensing position and companies that want to be recommended by AI assistants land in different places, and both are legitimate.
Who maintains this methodology
MaxAEO works on AI answer-engine visibility. The prompt set, crawler matrix and citability checklist above are the same ones our own product runs, published here and free to read without an account. The same methodology is packaged as a ChatGPT plugin and as open-source skills for Claude and Codex, so you can run it inside your own assistant.
A single run is only a snapshot. Turning it into numbers you can compare means freezing the prompt set, re-running it across engines on a fixed schedule and keeping the history — the part that is most tedious by hand, and the part our product automates.