LLM Visibility Score Formula: A Reproducible Measurement Framework

by

·

LLM Visibility Score Formula: A Reproducible Measurement Framework

By maxaeo.ai | Published 2026-09-26 | Updated 2026-09-26

An LLM visibility score formula should measure how often a brand appears in relevant AI answers and how prominently it is presented. A defensible score uses a fixed prompt set, multiple engines, repeated observations, explicit position weights, and an auditable denominator—not an unexplained number from a single query.

What Is an LLM Visibility Score?

An LLM visibility score is a normalized estimate of a brand’s exposure across a defined sample of AI-generated answers. It measures observable outputs from systems such as ChatGPT, Gemini, Perplexity, or Claude; it does not reveal or reverse-engineer their proprietary ranking algorithms.

The simplest version is mention rate:

Mention Rate = Responses Mentioning the Brand ÷ Valid Responses × 100

If a brand appears in 30 of 100 valid responses, its mention rate is 30%. This is transparent, but it treats a first recommendation and a minor mention near the end as equally valuable.

Measurement must also account for variation. A September 2026 study covering 50 questions, six engines, and 15 repeated runs found that one run captured only 62%–77% of the brands observed across five runs. This supports repeated sampling rather than one-off checks (Repeated Queries Exhaust an LLM’s Brand Recommendations but Not Its Sources). (arxiv.org)

The Recommended Visibility Formula

A practical formula should combine presence and prominence, while leaving sentiment, citations, and competitive share as separate diagnostic metrics. This preserves interpretability: the score answers “How visible are we?” without mixing exposure with reputation or source quality.

For each valid response (i), calculate:

Response Visibilityᵢ = Mᵢ × [0.60 + 0.40 ÷ log₂(Pᵢ + 1)]

Then aggregate the responses:

LLM Visibility Score =
100 × Σ(Wᵢ × Response Visibilityᵢ) ÷ Σ(Wᵢ)

Where:

Variable Meaning
Mᵢ 1 when the brand is mentioned; otherwise 0
Pᵢ The brand’s ordinal position among named alternatives
Wᵢ A predetermined prompt, engine, or market weight
0.60 Presence floor for any qualifying mention
0.40 Additional credit for appearing prominently

The 60/40 split is a proposed governance default, not an industry standard. It prevents lower-position mentions from becoming nearly worthless while still rewarding brands presented first.

LLM visibility score formula showing mention, position, and observation weights

How Do You Calculate the Score?

Calculate the score through a controlled observation process. Keep the methodology unchanged between reporting periods so that movement reflects changing AI answers rather than altered prompts, engines, or scoring rules.

  1. Define buyer prompts. Include unbranded discovery, category comparison, use-case, alternative, and purchase-intent questions. A prompt gap analysis can identify missing buyer journeys.
  2. Choose engines and markets. Record the platform, language, country, model mode, and whether web retrieval is active.
  3. Repeat important prompts. Use independent sessions and more than one run when budgets permit.
  4. Save every raw answer. Preserve the answer, cited sources, timestamp, mention status, position, and recommendation language.
  5. Apply fixed coding rules. Decide beforehand how aliases, product names, spelling variants, refusals, and failed responses will be handled.
  6. Calculate response-level values. Apply the presence and position formula to each valid observation.
  7. Aggregate with disclosed weights. Use equal weights by default. Apply intent weights only when they are defined before collection.
  8. Report uncertainty. Show the sample size, run-to-run range, or a bootstrap confidence interval beside the score.

Worked Example: Same Mention Rate, Different Visibility

Consider an illustrative test containing three buyer prompts, two AI engines, and two independent runs per prompt. This produces 12 equally weighted observations.

Brand A appears seven times at positions 1, 2, 2, 3, 4, 1, and 5. Its response values total approximately 6.03:

Mention Rate = 7 ÷ 12 × 100 = 58.3%

LLM Visibility Score = 6.03 ÷ 12 × 100 = 50.3

Brand B also appears seven times, but at positions 3, 3, 3, 4, 4, 5, and 5. Its response values total approximately 5.45:

Mention Rate = 7 ÷ 12 × 100 = 58.3%

LLM Visibility Score = 5.45 ÷ 12 × 100 = 45.5

Both brands have identical mention frequency, yet Brand A earns a higher score because it appears earlier. This small original example demonstrates why mention rate alone can conceal meaningful differences in recommendation prominence.

For relative competitive exposure, calculate share of voice in LLM responses separately.

Which Metrics Should Stay Outside the Formula?

Sentiment, citations, factual accuracy, and recommendation strength should accompany the visibility score rather than be hidden inside it. Combining them can produce misleading results—for example, a frequently criticized brand could receive a high composite score simply because negative mentions increased.

Use a scorecard with distinct fields:

Metric Question Answered
Visibility score How often and how prominently does the brand appear?
Recommendation rate How often is the brand explicitly suggested?
Citation rate How often is the brand’s domain used as a source?
Citation share Which external domains influence the answers?
Sentiment Is the brand described positively, neutrally, or negatively?
Factual accuracy Are product claims and descriptions correct?
Competitive share of voice How much category exposure belongs to the brand?

Citation count also differs from citation influence. Research analyzing 602 prompts and more than 21,000 valid search-layer citations found that source breadth and the degree to which a source shapes an answer can diverge (Citation Selection to Citation Absorption). (arxiv.org)

How Should the Score Be Reported?

A credible report pairs the headline score with an evidence card. At minimum, disclose the prompt count, eligible-response denominator, engines, markets, languages, repetitions, collection dates, weighting rules, formula version, and missing-response policy.

Dashboard evidence card accompanying an LLM visibility score formula

Avoid universal labels such as “good” or “bad.” A score depends on category breadth, prompt difficulty, market, engine mix, and competitor set. Compare it with:

  • The same brand under an unchanged methodology over time
  • Named competitors measured with the same prompts
  • Engine-level and intent-level segments
  • Mention, recommendation, citation, and sentiment trends

MaxAEO monitors brand mentions, citations, recommendations, sentiment, competitive position, and source patterns across eight AI engines with daily updates. Teams can use its AI citation measurement framework to design supporting metrics or generate a free AI visibility diagnostic from MaxAEO.

Frequently Asked Questions

Is there one standard LLM visibility score formula?

No. The industry has not adopted a universal formula. Some systems report simple mention rate, while others add position, citations, sentiment, consistency, or source authority. Scores from different methodologies should not be compared directly.

How many prompts are needed?

There is no universal minimum. Use enough prompts to represent distinct buyer intents, then favor repeated runs over a large collection of near-duplicate questions. Always disclose the prompt and observation counts.

Should branded prompts be included?

Branded prompts are useful for testing description accuracy and reputation, but they inflate discovery visibility. Report branded and unbranded prompts as separate segments rather than averaging them without disclosure.

Should failed AI responses remain in the denominator?

Define the rule before collection. Technical failures are usually excluded from valid responses, while genuine refusals may remain if refusal behavior is relevant to the prompt. Apply the policy consistently across brands and engines.

How often should visibility be measured?

Frequent monitoring helps reveal changes, but individual daily movements can be noisy. Use daily collection when available, then evaluate rolling trends and consecutive periods before attributing a change to content, PR, or technical work.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →