By maxaeo
AI visibility score calculation converts a fixed set of AI responses into a 0–100 measure of how often and how strongly a brand appears. A reproducible score defines the observation unit, coding rules, weights, denominator, missing-data policy, and uncertainty before collection, then preserves every raw answer for audit.
There is no universal industry formula. Vendors may measure mentions, recommendations, rank, citations, sentiment, or share of voice—and label each result an “AI visibility score.” Absolute scores are therefore comparable only when the underlying methods match.
This guide presents Open Scorecard v1.1, a proposed maxaeo framework. Its weights and quality thresholds are governance defaults, not external standards. The framework’s value is that another analyst can rebuild the result from the same response-level data.
For the broader metric definition and alternative models, see AI Visibility Score: Definition, Formula, and Scorecard.
How Do You Calculate an AI Visibility Score?
Score every valid AI response for mention, recommendation, prominence, citation support, description accuracy, and competitive share. Multiply those values by disclosed component weights, convert the result to 0–100, and calculate the weighted average across prompts, platforms, markets, and runs.
The calculation has eight steps:
- Define the scope: brands, competitors, prompts, platforms, markets, languages, and reporting period.
- Freeze the prompt library: do not add favorable prompts after seeing results.
- Collect the eligible observation grid: one response for every planned prompt × platform × market × run.
- Code each response: apply the same written rules to every brand.
- Calculate a row score: combine the six component values.
- Apply observation weights: use documented prompt, platform, market, and time priorities.
- Report measurement quality: publish coverage, uncertainty, reviewer agreement, and sensitivity.
- Version the method: preserve the formula, prompt set, coding guide, and raw answers used for every reported score.
The aggregate formula is:
[
\text{AI Visibility Score} =
\frac{\sum_{i=1}^{n} w_i \times \text{Row Score}i}
{\sum{i=1}^{n} w_i}
]
Only valid observations belong in this point estimate. Missing observations require separate coverage reporting and bounds.
Which AI Visibility Metric Do You Actually Need?
Mention rate, recommendation rate, citation rate, share of voice, and a composite visibility score answer different questions. Do not use one label for all five.
| Metric | Basic calculation | Question answered |
|---|---|---|
| Mention rate | Responses naming the brand ÷ valid responses | How often does the brand appear? |
| Recommendation rate | Responses endorsing or shortlisting the brand ÷ valid responses | How often is the brand recommended? |
| Citation rate | Responses with a supporting citation for the brand ÷ valid responses | How often is the brand’s presence supported by a visible source? |
| AI share of voice | Weighted brand presence ÷ weighted presence of all tracked brands | How much competitive attention does the brand receive? |
| Composite visibility score | Weighted combination of multiple coded outcomes | How broad and strong is the brand’s overall presence? |
Use mention rate when simplicity and portability matter most. Use a composite when stakeholders need to distinguish recognition from recommendation, prominence, evidence, and accuracy.
A composite provides more diagnostic value, but every added component introduces another assumption that must be documented and tested.
What Does Open Scorecard v1.1 Measure?
Open Scorecard uses six response-level components:
| Component | Symbol | Coding rule | Default weight |
|---|---|---|---|
| Mention | (M) | 1 if the answer body names the tracked brand; otherwise 0 | 30% |
| Recommendation | (R) | 1 if the answer explicitly endorses the brand or includes it in a responsive shortlist; otherwise 0 | 25% |
| Prominence | (P) | 1.00 for first or sole choice; 0.75 for positions 2–3; 0.50 for another or unranked shortlist position; 0.25 for a narrative mention; 0 if absent | 15% |
| Citation support | (C) | 1.00 when a visible source directly supports a material brand claim; 0.50 when relevant but only partially supportive; 0 otherwise | 15% |
| Description quality | (Q) | Mention multiplied by an accuracy grade of 1.00, 0.50, or 0 | 10% |
| Competitive share | (V) | (M/k), where (k) is the number of tracked brands named in the answer; 0 when the tracked brand is absent | 5% |
The row formula is:
[
\text{Row Score}_i =
100 \times
(0.30M_i + 0.25R_i + 0.15P_i + 0.15C_i + 0.10Q_i + 0.05V_i)
]

The default model gives 55% of the score to reach and recommendation. The other 45% measures the quality and competitive strength of that presence. These weights are deliberately editable because a publisher, SaaS company, retailer, and local service business may value different outcomes.
Coding Rules That Prevent Ambiguous Scores
A calculation is only reproducible when edge cases are resolved in advance:
- Do not count the prompt itself. A brand named by the user but absent from the generated answer has (M=0).
- Use an approved alias list. Count verified product names, abbreviations, and former names only when they clearly identify the tracked entity.
- Do not count a source card as an answer-body mention. It may qualify for (C), but not (M), unless the brand also appears in the generated response.
- Distinguish a recommendation from an example. “Consider Brand X” or inclusion in a requested shortlist earns (R=1). A historical or incidental reference earns (R=0).
- Evaluate visible citations only. Do not infer which documents a model may have used internally.
- Grade description quality against a fact sheet. Use (Q=1) when all material audited claims are correct, (Q=0.5) for a minor error or a mention without a checkable description, and (Q=0) for absence, misidentification, or a material error.
- Cap each component at 1. Multiple mentions or citations do not push a response above the component maximum.
Store the coding guide and approved fact sheet beside the dataset. Otherwise, later reviewers cannot reproduce accuracy and recommendation judgments.
What Is the Correct Observation Unit?
The observation unit is one prompt × platform × market × run. Every row must represent one independently captured AI response with its collection status and raw evidence.
For example:
- 100 prompts
- 4 platforms
- 2 markets
- 3 scheduled runs
This design creates (100 \times 4 \times 2 \times 3 = 2{,}400) expected observations.
Repeated runs of one prompt are not equivalent to new prompts. They measure answer volatility, while additional prompts increase coverage of user needs and decision contexts. A hundred runs of one question do not provide the same evidence as 100 distinct questions.
How Should Prompts and Platforms Be Weighted?
Use equal weights unless documented business evidence supports a different choice. Set all weights before collection, and retain an equal-weight benchmark so stakeholders can see how much the business weighting changes the result.
The observation weight is:
[
w_i =
\text{Prompt Weight} \times
\text{Platform Weight} \times
\text{Market Weight} \times
\text{Recency Weight}
]
Prompt Weights
A practical starting scale is:
- 3 — Decision prompts: category, comparison, shortlist, or purchase-selection questions
- 2 — Consideration prompts: use case, problem, integration, workflow, or alternative questions
- 1 — Educational prompts: definitions, explanations, and general research questions
This is a business-priority scale, not estimated search volume. Replace it with documented pipeline, customer-research, or market data when reliable inputs exist.
Report branded and non-branded prompts separately. If branded prompts contribute to the headline score, Open Scorecard caps them at 10% of expected prompt weight because the brand is already supplied in the question. The 10% cap is a proposed anti-inflation control, not an industry standard.
Platform and Market Weights
Equal platform weights are the safest default. Change them only when customer research or first-party attribution shows that a platform materially matters more to the target audience.
Market weights can reflect revenue, pipeline, customer count, or strategic priority. State which basis was used; do not describe a strategic allocation as audience market share.
Recency Weights
Use equal time weights inside a fixed reporting period. For a rolling score, disclose the decay rule. One reproducible option is:
[
\text{Recency Weight} =
2^{-\text{Age in Days}/\text{Half-Life in Days}}
]
A 28-day half-life gives a response collected 28 days ago half the weight of a response collected today. Keep a fixed-window benchmark alongside the rolling score because decay can make the metric move even when no answer changes.
How Do You Collect Reproducible Inputs?
A reproducible dataset preserves the prompts, settings, timestamps, raw answers, source links, collection failures, labels, and weights used in the score. Saving only the final labels or dashboard summary is insufficient.
Use this workflow:
- Define tracked brands, canonical aliases, competitors, markets, languages, and eligible platforms.
- Build prompts across category, comparison, alternative, problem, use-case, integration, educational, and branded intent.
- Assign intent labels and weights before collecting responses.
- Freeze prompt text and create a versioned prompt-set identifier.
- Run the same eligible grid on a documented schedule.
- Use isolated conversations where possible so earlier prompts do not influence later answers.
- Record interface, model or product label, market settings, timestamp, and any personalization state that can be observed.
- Save the complete response, visible citations, and collection status.
- Apply the written coding guide.
- Review ambiguous rows and close the reporting window before calculating the score.
Four controls prevent artificial score gains:
- A non-branded prompt cannot contain the tracked brand or a unique product cue.
- Failed collections cannot silently disappear from the denominator.
- New prompts cannot be added only because the brand performs well on them.
- A formula or coding change creates a new methodology version and a break in the trend line.
A compact methodology fingerprint might read:
score=v1.1; prompt_set=2026-Q3-01; coding=1.0;
weights=30/25/15/15/10/5; scope=4-platforms/2-markets;
branded_cap=10%; missing_policy=exclude-and-bound
Publishing this fingerprint beside the score makes silent methodology drift easier to detect.
Worked AI Visibility Score Calculation
This six-row teaching example produces a score of 63.0, weighted coverage of 91.7%, and a full-coverage missing-data range of 57.7–66.0. It demonstrates the arithmetic only; it is too small to serve as a platform or industry benchmark.
Platform, market, and recency weights are all 1 in this example:
| Prompt intent / platform | Prompt weight | Status | (M) | (R) | (P) | (C) | (Q) | (V) | Row score |
|---|---|---|---|---|---|---|---|---|---|
| Category / ChatGPT | 3 | Valid | 1 | 1 | 1.00 | 0.50 | 1.00 | 0.50 | 90.0 |
| Category / Gemini | 3 | Valid | 1 | 1 | 0.75 | 1.00 | 1.00 | 0.33 | 92.9 |
| Use case / ChatGPT | 2 | Valid | 1 | 0 | 0.25 | 0 | 0.50 | 0.25 | 40.0 |
| Use case / Gemini | 2 | Valid, absent | 0 | 0 | 0 | 0 | 0 | 0 | 0.0 |
| Definition / ChatGPT | 1 | Valid | 1 | 0 | 0.25 | 1.00 | 1.00 | 1.00 | 63.8 |
| Definition / Gemini | 1 | Timeout | — | — | — | — | — | — | Missing |
The valid weighted score is:
[
\frac{
(3 \times 90)
+(3 \times 92.9167)
+(2 \times 40)
+(2 \times 0)
+(1 \times 63.75)
}{
3+3+2+2+1
}
62.95
]
Rounded to one decimal place:
[
\boxed{\text{AI Visibility Score}=63.0}
]
Expected weight is 12, while valid weight is 11:
[
\text{Coverage}=\frac{11}{12}=91.7%
]
This example also shows why the composite must be accompanied by its components. The zero is a real observed absence; the timeout is missing evidence. Treating both as zero would answer a different question.
How Should Missing Data Be Handled?
A valid non-mention is zero. A technical failure is missing. Exclude genuine failures from the point estimate, publish weighted coverage, and calculate the lowest and highest possible full-coverage scores.
| Event | Treatment | Reason |
|---|---|---|
| Valid answer does not name the brand | Score all applicable components as 0 | Absence was observed |
| Valid answer refuses an eligible, on-topic request | Score as 0 | The refusal was the user-visible result |
| Timeout, API error, or corrupted response | Mark missing | No complete answer was observed |
| Answer displays no source links | Set citation support to 0 | Lack of a visible citation is observable |
| Source capture fails while citations are part of the score | Mark the observation missing | Required evidence is incomplete |
| Accuracy is difficult to adjudicate | Send for review | Label uncertainty is not technical missingness |
| Retry creates duplicate captures | Keep the first complete planned response | Prevents accidental extra weight |
| Platform was known to be ineligible before collection | Remove it from the expected grid | Ineligibility must not be decided after seeing results |
Weighted coverage is:
[
\text{Coverage} =
\frac{\sum w_{\text{valid}}}
{\sum w_{\text{expected eligible}}}
]
Open Scorecard uses these predeclared reporting gates:
- 90% or higher: publish with coverage shown
- 80%–89.9%: label the composite provisional
- Below 80%: withhold the headline composite and report component or platform results
These are maxaeo governance defaults, not statistical standards.
Worst-Case and Best-Case Bounds
If missing observations had all scored 0:
[
\text{Worst Case} =
\text{Point Score} \times \text{Coverage}
]
If every missing observation had scored 100:
[
\text{Best Case} =
(\text{Point Score} \times \text{Coverage})
+100(1-\text{Coverage})
]
For the worked example:
[
\text{Worst Case} =
63.0 \times 0.9167
=57.7
]
[
\text{Best Case} =
(63.0 \times 0.9167)
+(100 \times 0.0833)
=66.0
]
The defensible full-coverage range is therefore 57.7–66.0. This range describes uncertainty caused by missing data, not sampling uncertainty.
How Do You Measure Statistical Uncertainty?
Resample prompt groups rather than individual response rows. Platforms and repeated runs for the same prompt are related observations, so treating every row as independent usually produces confidence intervals that are too narrow.
Use a prompt-cluster bootstrap:
- Group every platform, market, and run by
prompt_id. - If the library is stratified, keep separate groups for intent and market.
- Sample prompt IDs with replacement within each stratum.
- Bring all associated response rows into the resampled dataset.
- Recalculate the complete weighted score.
- Repeat 2,000 times.
- Use the 2.5th and 97.5th percentiles as the 95% interval.
When repeated-run volatility is a major concern, use a two-stage bootstrap: sample prompt groups first, then sample runs within each selected prompt.
A confidence interval does not predict every future answer. It estimates how the score varies under repeated sampling from the tracked prompt design.
Is a Score Change Statistically Meaningful?
For period-over-period reporting, use a paired bootstrap when both periods contain the same prompts:
- Resample prompt IDs.
- Calculate the score for period A and period B using the same sampled prompts.
- Store the difference (B-A).
- Repeat 2,000 times.
- Report the 95% interval for the change.
If the interval includes zero, describe the movement as directional rather than confirmed. If the prompt library changed, report both:
- A matched score using prompts present in both periods
- An expanded score using the complete new library
This disclosure-first treatment is consistent with the NIST AI RMF: Generative Artificial Intelligence Profile, which emphasizes documented measurement, limitations, and ongoing monitoring of generative AI behavior.
How Should Coding Reliability Be Checked?
Double-code a sample before publishing the composite. A mathematically precise score is not reliable when reviewers disagree about what qualifies as a recommendation, supporting citation, or material accuracy error.
Open Scorecard’s operating procedure is:
- Double-code at least 10% of rows or 100 rows, whichever is larger, capped at the dataset size.
- Report percentage agreement for every component.
- Use Cohen’s kappa for binary labels when there are two reviewers.
- Use weighted kappa for prominence and other ordinal grades.
- Treat 0.80 as the predeclared operating threshold for agreement—not a universal law.
- If a component falls below the threshold, clarify the guide and recode affected rows.
Do not resolve disagreements by averaging incompatible labels. Adjudicate them against the written rule and record the final reviewer.
Which Sensitivity Tests Are Required?
Recalculate the score under reasonable alternative assumptions. If a small change in weights, prompt mix, or platform coverage causes a large movement, the score is method-dependent and should be labeled accordingly.
At minimum, test:
- Equal component weights
- Equal prompt weights
- Stricter recommendation coding
- Stricter citation-support coding
- Branded prompts excluded
- Each platform removed in turn
- Each major prompt-intent group removed in turn
- Business-weighted versus equal-weight platform mix
Using the six-row teaching sample:
| Sensitivity test | Score | Change from 63.0 |
|---|---|---|
| Default method | 63.0 | — |
| Equal component weights | 58.3 | -4.7 |
| Equal prompt weights | 57.3 | -5.7 |
| Partial citations changed from 0.50 to 0 | 60.9 | -2.1 |
| Strongest platform removed | 55.8 | -7.2 |
The sample is highly sensitive, which is expected from only three prompts. It should not be used for decisions without expanding the prompt library.
Open Scorecard uses these interpretation bands:
- Stable: every reasonable test moves the score by less than 2 points
- Moderately sensitive: at least one test moves it by 2–5 points
- Highly sensitive: at least one test moves it by more than 5 points
These bands are governance thresholds. Their purpose is consistent disclosure, not a claim that five points has universal statistical meaning.

How Can You Reproduce the Score in a Spreadsheet?
Export one row per planned observation, including failures. The spreadsheet must contain enough evidence to rebuild component values, row weights, coverage, and the final score without relying on a dashboard transformation.
Recommended columns are:
period
prompt_id
prompt_text
prompt_intent
prompt_weight
platform
platform_weight
market
market_weight
run_id
recency_weight
collected_at
valid_observation
missing_reason
brand_mentioned
brand_recommended
prominence_value
citation_value
accuracy_grade
tracked_brands_mentioned
competitive_share
row_score
response_text
citation_urls
reviewer
coding_version
prompt_set_version
score_version
Calculate competitive share as:
IF(brand_mentioned=1, 1/tracked_brands_mentioned, 0)
Calculate each valid row score as:
=100*(0.30*M2+0.25*R2+0.15*P2+0.15*C2+0.10*Q2+0.05*V2)
Calculate observation weight as:
=PromptWeight*PlatformWeight*MarketWeight*RecencyWeight
With Valid coded as 1 or 0, calculate the aggregate as:
=SUMPRODUCT(RowScoreRange,WeightRange,ValidRange)
/SUMPRODUCT(WeightRange,ValidRange)
Calculate coverage as:
=SUMPRODUCT(WeightRange,ValidRange)/SUM(WeightRange)
The spreadsheet result should match the dashboard within the published rounding tolerance. If it does not, the system is applying an undocumented filter, transformation, or normalization.
How Many Prompts Are Needed?
There is no universal minimum. Prompt diversity, segment coverage, volatility, and the required decision precision matter more than an arbitrary row count. Expand the library until commercially important intents are represented and the prompt-cluster interval is narrow enough for the decision.
A defensible prompt library should cover:
- Category discovery
- Direct comparisons
- Alternatives and replacements
- Problems and desired outcomes
- Industry or persona-specific use cases
- Integrations and constraints
- Educational definitions
- Branded verification prompts, reported separately
Repeated runs cannot compensate for missing intent coverage. If removing one small prompt group changes the composite materially, rebalance or expand the library before treating the score as stable.
How Should the Score Be Interpreted?
The composite identifies where to investigate; the component pattern identifies what may need to change. A three-point gain from branded mentions is not equivalent to a three-point gain in non-branded recommendations.
| Component pattern | Likely interpretation |
|---|---|
| High mention, low recommendation | The brand is recognized but not shortlisted |
| High recommendation, low citation support | The brand is endorsed, but visible evidence is weak or inconsistent |
| High citation support, low mention | Useful brand content is cited in a narrow set of answers |
| Low description quality | AI answers contain outdated, incomplete, or incorrect positioning |
| High competitive share, low mention | The brand performs well when present but appears in too few responses |
| Branded score rising, non-branded score flat | Awareness within known-brand prompts improved, but discovery did not |
A useful reporting view includes:
- Composite score, confidence interval, and coverage
- Mention and recommendation rates
- Citation-support rate and cited domains
- Description errors by claim
- Branded versus non-branded performance
- Intent, market, and platform breakdowns
- New, lost, and changed recommendations
- Sensitivity status
- Prompt-set and methodology versions
The AI visibility dashboard metrics marketing teams should monitor weekly explains how to connect the headline score to these diagnostic measures.
Once gaps are identified, use the AI visibility prioritization framework to rank fixes by business importance, evidence gap, expected impact, and implementation effort.
Does a Higher Score Prove GEO Performance?
No. A higher score shows improved performance inside the declared measurement design. It does not prove increased revenue, universal AI visibility, or causation by a specific generative engine optimization change.
Validate commercial impact separately through:
- AI referral sessions and assisted conversions
- Qualified demo or signup attribution
- Sales-call mentions of AI discovery
- Branded search changes
- Recommendation gains among target customer segments
- Share of commercially important shortlist prompts
Maintain an intervention log for content releases, product-positioning changes, structured data, third-party coverage, and major site updates. Compare affected prompts with unaffected prompts over the same period.
Models, retrieval indexes, source selection, interfaces, and personalization can change during an evaluation window. Report the evidence and timing without claiming guaranteed causation.
How Should AI Visibility Tools Be Compared?
Compare methodologies before comparing headline scores. Two vendors’ absolute numbers are interchangeable only when their prompts, platforms, markets, coding rules, weights, collection periods, and missing-data policies match.
Ask each vendor:
- Can the score be rebuilt from exported response-level observations?
- Is the exact formula published?
- Are component, prompt, platform, and market weights visible?
- Is a valid non-mention distinguished from a failed collection?
- Are raw answers and visible citations retained?
- Can users inspect recommendation and accuracy labels?
- Are branded prompts separated from non-branded prompts?
- Are historical scores recalculated after formula changes?
- Is the methodology version displayed for every reporting period?
- Are coverage, sample size, confidence limits, and sensitivity disclosed?
- Can the prompt library and row-level data be exported?
- How long are complete responses and citations retained?
Pricing can also hide material methodological differences. The AI visibility tool pricing comparison guide explains how to compare prompt limits, platform coverage, data retention, and export access rather than relying on a monthly price alone.
If a vendor cannot disclose or export enough information to reproduce its score, treat the result as a proprietary trend index. It may still be useful inside that product, but it is not a portable cross-vendor benchmark.
What Are the Formula’s Main Limitations?
Transparency makes the score auditable, not universally correct. Prompt selection, label definitions, component overlap, platform behavior, and business weights still influence the result.
The main limitations are:
- Component overlap: mention, recommendation, and prominence are correlated because a recommended brand must normally be named.
- Platform asymmetry: some interfaces expose citations, sources, locations, or personalization more consistently than others.
- Prompt-frame dependence: changing the proportion of comparison, educational, or branded prompts can change the score without changing any individual answer.
- Human judgment: recommendation, citation support, and accuracy grades require written rules and reviewer checks.
- Temporal volatility: a response is evidence from one collection event, not a permanent ranking.
- No universal benchmark: a score of 60 has meaning only relative to the same method, competitors, scope, and historical periods.
Do not silently revise the method. Preserve old exports and show a break in the trend line after a material formula, scope, or coding change.
What Must a Publishable Scorecard Include?
A publishable scorecard needs eight visible elements: formula, component results, prompt and platform scope, coverage, missing-data bounds, confidence interval, sensitivity result, and methodology version. A score without this context cannot be independently interpreted.
A production statement should follow this pattern:
Open Scorecard v1.1: [score]/100, based on [prompt count] prompts across [platform count] platforms and [market count] markets during [period]; weighted coverage [percentage]; missing-data range [low–high]; 95% prompt-cluster interval [low–high]; [stable/moderately sensitive/highly sensitive]; prompt set [version].
For the small teaching dataset in this article, the honest statement is:
Open Scorecard v1.1: 63.0/100, based on three prompts and two platforms; weighted coverage 91.7%; missing-data range 57.7–66.0; no confidence interval reported because prompt diversity is insufficient; highly sensitive, with a maximum tested movement of 7.2 points.
That statement is more useful than an isolated dashboard number because it reveals what was measured, how much evidence was captured, and how strongly the result depends on methodological choices.
Frequently Asked Questions
Is There a Standard AI Visibility Score Formula?
No. AI visibility vendors use different prompt sets, platforms, denominators, component definitions, and weights. A formula can be transparent and useful without being an industry standard. Absolute scores should not be compared unless the methods and input data match.
What Is a Good AI Visibility Score?
There is no universal good-score threshold. Evaluate performance against the brand’s historical baseline and declared competitors using the same methodology. Prioritize commercially important non-branded prompts, component improvement, coverage, confidence, and sensitivity rather than an arbitrary 0–100 grade.
How Often Should AI Visibility Be Calculated?
Collect often enough to measure answer volatility and report on a schedule suited to the decision. Daily or multi-day collection can support monitoring, while weekly reporting reduces noise for marketing teams. Executive reporting should include longer trends, coverage, uncertainty, and explanations for material changes.
Is Mention Rate the Same as an AI Visibility Score?
No. Mention rate is the percentage of valid answers that name the brand. A composite score may also include recommendations, prominence, citations, description accuracy, and competitive share. Mention rate is simpler and more portable; a composite is more diagnostic but depends on more assumptions.
How Many Prompts Are Required?
There is no universal minimum. The library must represent the decisions, problems, personas, use cases, markets, and languages that matter to the business. Use prompt-cluster confidence intervals and sensitivity tests to determine whether the sample is stable enough for the intended decision.
Can Scores from Two AI Visibility Tools Be Compared?
Only when both tools use the same prompts, platforms, markets, collection period, coding definitions, weights, and missing-data policy. Otherwise, compare trends within each tool rather than treating their absolute scores as equivalent. Response-level exports are necessary for a true reconciliation.