Brand Co-Mention Analysis: Method, Metrics, and Example

by

·

Brand co-mention analysis network showing substitute, complement, category-leader, and challenger relationships

Brand co-mention analysis measures how often brands appear together in the same AI-generated answer and interprets the surrounding language, position, prompt intent, and change over time. It reveals whether answer engines frame brands as substitutes, complements, category leaders, or emerging challengers—and how reliable each relationship is.

The result is an AI-mediated consideration set: the companies and products that answer engines place around your brand when responding to relevant questions. It is not proof of what buyers ultimately consider, but it can expose category associations that traditional keyword rankings, social listening, and sales-defined competitor lists miss.

A defensible analysis follows eight steps:

  1. Define the buyer decisions you want to study.
  2. Build an intent-balanced prompt set.
  3. Collect complete answers under documented conditions.
  4. Resolve brand names, products, and aliases.
  5. Measure co-mentions at answer, passage, and sentence level.
  6. Label the relationship expressed in context.
  7. calculate support, affinity, direction, and recurrence.
  8. Validate important edges across prompts, engines, and dates.

What does brand co-mention analysis measure?

It measures both the frequency and meaning of brand pairings within a defined set of generated answers. A raw count identifies which brands appear together. A complete analysis determines why they appear together, which brand receives more prominence, and whether the association is stable.

Suppose Brand A and Brand B appear together in 40 answers. That number is incomplete without knowing:

  • Whether the products were presented as alternatives or integrations
  • Whether one brand consistently appeared first
  • Which prompt intents generated the pairing
  • Whether the relationship appeared across multiple answer engines
  • Whether the same few prompt variants produced most occurrences
  • Whether the association strengthened or weakened over time
  • Which cited sources supported the description
Analysis method Primary unit Best for Main limitation
Web co-occurrence Page or domain Finding broad entity associations Does not show how AI systems assemble answers
Social listening Post or conversation Measuring discussion volume and sentiment Rarely captures recommendation order or buyer intent
Keyword overlap Search query Finding domains competing in organic results Does not identify direct answer-level relationships
AI share of voice Generated answer Measuring visibility within a prompt set Does not explain which brands surround yours
Brand co-mention analysis Answer and passage Mapping roles, shortlists, and relationship changes Cannot establish causation from outputs alone
Brand co-mention analysis network showing substitute, complement, category-leader, and challenger relationships

Why do AI co-mentions matter?

Repeated co-mentions reveal how answer engines organize a market for a particular question. That organization may differ from your internal category definition because AI answers combine learned entity associations, retrieved sources, prompt wording, location, product availability, and user context.

A direct commercial competitor may appear rarely if answer engines cannot find strong evidence about it. An open-source project, consultancy, or adjacent platform may appear frequently because it offers another route to the same buyer outcome.

For example:

  • “Best enterprise observability platforms” is likely to produce substitutes.
  • “How should a startup monitor Kubernetes?” may combine platforms, cloud services, and open-source tools.
  • “Tools that integrate with Acme” should produce complements.
  • “Alternatives to Atlas for regulated teams” may reveal challengers within a specific segment.

Pooling those answers into one competitor count would erase the relationships that make the data actionable.

Use co-mention analysis alongside customer interviews, win-loss data, CRM records, and search demand. The network shows how AI answers frame the market, not the entire market itself.

Which questions can the analysis answer?

A well-designed study can answer six practical questions:

  1. Who enters the shortlist with us?
    Identify brands that repeatedly appear in the same recommendation answers.

  2. Why are those brands present?
    Distinguish substitutes, complements, category anchors, and emerging challengers.

  3. Where is our positioning unclear?
    Find prompt clusters in which the brand is assigned the wrong category, audience, price tier, or use case.

  4. Who controls the category narrative?
    Detect brands that appear first, define terminology, or serve as the reference point for comparisons.

  5. Which relationships are changing?
    Track brands entering new prompt clusters or moving into stronger recommendation positions.

  6. What evidence shapes the answers?
    Review the pages and citations supporting important pairings, descriptions, and rankings.

How to conduct brand co-mention analysis

1. Define the decision scope

Start with a specific market question, not a list of brand names. “Map every brand related to analytics” is too broad to produce a useful network. “Map which platforms are recommended to mid-market SaaS teams evaluating product analytics” is measurable.

Document:

  • Product or business unit
  • Audience and company size
  • Geography and language
  • Use cases
  • Buying stage
  • Relevant answer engines
  • Analysis period
  • Decisions the findings will inform

The scope determines which prompts and co-mentions count as relevant.

2. Build an intent-balanced prompt set

Use prompt clusters that represent different stages and types of buyer inquiry. A prompt set dominated by “alternatives” will manufacture a substitute-heavy network even if buyers ask broader questions.

Include these clusters where relevant:

Prompt cluster Example Relationship most likely to surface
Category discovery “What are the best tools for X?” Leaders and substitutes
Alternatives “What are alternatives to Brand A?” Substitutes and challengers
Direct comparison “Brand A vs Brand B for Y” Substitutes and segment differences
Use case “What should a team use to accomplish X?” Mixed consideration set
Pain point “How can a company solve X?” Cross-category options
Integration “What works with Brand A?” Complements
Migration “What should replace Brand A?” Substitutes
Audience-specific “Best X for regulated enterprises” Segment leaders and challengers

Create controlled wording variants, but group them under a stable prompt-cluster ID. Give each cluster equal or intentionally documented weight so near-duplicate variants cannot dominate the results.

See the full method for creating an AI brand-monitoring prompt set.

3. Collect complete answers under controlled conditions

Store the full answer and its collection context, not just extracted brand names. Without the underlying evidence, another analyst cannot verify the relationship or determine why an edge exists.

The minimum useful record contains:

Field Purpose
answer_id Unique record identifier
prompt_id and cluster_id Separates individual prompts from buyer intents
prompt_text Preserves the exact input
engine and surface Identifies where the answer appeared
model_or_version Records the version when disclosed
locale Preserves geographic and language context
collected_at Enables time-series analysis
answer_text Supports review and reprocessing
brand_entities Stores normalized brands detected
mention_positions Captures order and prominence
citations Identifies supporting sources
screenshot_or_snapshot Preserves evidence if the answer changes

Record whether the session was logged in, personalized, or influenced by prior conversation. If those conditions cannot be controlled, disclose them as limitations rather than mixing the results silently.

4. Resolve entities before counting

Normalize aliases while keeping distinct commercial entities separate. Otherwise, spelling differences split one brand into several nodes, while parent companies and products are incorrectly merged.

An entity dictionary should contain:

  • Canonical brand name
  • Common abbreviations
  • Product names
  • Previous names
  • Parent company
  • Subsidiaries
  • Domains
  • Disambiguation notes

For example, a company and its flagship product may need separate nodes if buyers evaluate the product independently. Merge them only when the analysis question treats them as the same decision entity.

Review ambiguous names manually. A short brand name may also be an ordinary word, technical acronym, or unrelated company.

5. Measure more than one co-mention window

Use answer-level co-mentions for consideration-set discovery and narrower windows for relationship interpretation.

Window Definition Best use
Answer level Both brands appear anywhere in one answer Broad consideration-set mapping
Section or list level Both appear in the same answer section or shortlist Recommendation and comparison analysis
Paragraph level Both appear in one paragraph Stronger contextual association
Sentence level Both appear in one sentence Relationship labeling

Store each scope rather than choosing one permanently. An answer-level edge may be strategically important even when the brands never share a sentence, but it should receive less contextual confidence.

6. Label the relationship expressed in context

Do not classify every co-mentioned brand as a competitor. Review the sentence, passage, list heading, and prompt intent before assigning a role.

Role Typical language Commercial validation
Substitute “Alternative to,” “versus,” “instead of,” “similar option” Can buyers choose one instead of the other for the same job?
Complement “Integrates with,” “works alongside,” “used together” Can a customer reasonably buy or use both?
Category leader Defines the category, appears first, anchors comparisons Is the brand central across multiple prompt clusters?
Emerging challenger Enters new shortlists, gains position or prompt coverage Does the rise recur across dates or engines?
Shared attribute Both cited as examples of one feature, audience, or model Is the association narrower than the overall category?
Incidental Appears in an unrelated citation, disclaimer, or long list Would removing it change the decision context?

“A or B” usually signals substitution. “A with B” usually signals complementarity. Neither rule is sufficient without the surrounding buyer job.

Store a primary label, any secondary label, and a confidence value. The same pair may be substitutive in one prompt cluster and complementary in another.

The 4R framework for evaluating an edge

The 4R Co-Mention Framework evaluates each brand pair through Recurrence, Relevance, Relationship, and Rate of change. It prevents a large but ambiguous count from becoming a strategic conclusion.

Dimension Question Evidence
Recurrence Does the pair appear reliably? Joint-answer count, prompt clusters, engines, and collection dates
Relevance Does it appear in commercially meaningful questions? Audience, use case, buyer stage, and prompt weight
Relationship Why are the brands mentioned together? Language, list position, passage scope, and citations
Rate of change Is the association strengthening or weakening? Rolling support, prompt coverage, engine coverage, and rank

The four dimensions should remain visible separately. A single blended score can conceal an edge with high frequency but weak commercial relevance, or a small edge that is expanding rapidly in high-intent prompts.

Which metrics should you calculate?

Use absolute support and normalized affinity together. Raw counts favor widely mentioned brands, while ratios such as lift can exaggerate relationships based on only a few answers.

For brands (A) and (B) across (N) answers:

Metric Formula What it shows Main warning
Pair support (n(A \cap B)) Number of shared answers Favors common brands
Pair frequency (n(A \cap B) / N) Share of the dataset containing both Depends on prompt mix
Directional probability (P(B\mid A)=n(A \cap B)/n(A)) How often B appears when A appears Direction must be reported both ways
Jaccard similarity (n(A \cap B)/[n(A)+n(B)-n(A \cap B)]) Overlap relative to combined coverage Can be unstable at low counts
Lift ([n(A \cap B)/N]/[(n(A)/N)(n(B)/N)]) Whether the pair exceeds chance expectation High lift does not imply high support
Placement rate Shared answers where a brand ranks first divided by shared ranked answers Recommendation prominence Requires consistent position coding
Cluster coverage Prompt clusters containing the pair divided by eligible clusters Breadth of relevance Clusters must be defined before collection
Engine recurrence Engines containing the pair divided by engines tested Cross-engine stability Engines are not independent observations

Lift above 1 means the pairing occurred more often than expected from each brand’s individual frequency in the measured dataset. It does not prove that one brand caused the other to appear.

Separate edge strength from edge confidence

Strength describes how prominent the relationship is; confidence describes how much evidence supports the interpretation. Combining them into one number makes weak evidence look more certain than it is.

A practical dashboard can report:

  • Strength: pair support, Jaccard similarity, directional probability, and placement
  • Confidence: recurrence across dates, engines, prompt clusters, and human coding agreement
  • Role: substitute, complement, leader, challenger, shared attribute, or incidental
  • Trend: new, strengthening, stable, weakening, or disappeared

If one score is required for visualization, publish its formula and keep the component metrics available.

Weight prompt clusters, not prompt volume

When prompt clusters contain different numbers of variants, assign each cluster a total weight of 1 and divide that weight across its prompts. Otherwise, the cluster with the most editorial variants will create the strongest edges by construction.

For uncertainty estimates, resample prompt clusters rather than individual answers. Answers generated from variations of the same buyer question are correlated and should not be treated as fully independent observations.

When is a co-mention edge reliable?

An edge is decision-ready only when quantitative recurrence and contextual interpretation both pass review. There is no universal sample threshold, so teams should publish a policy appropriate to the risk of the decision.

The maxaeo method in this guide uses the following exploratory rule:

Evidence gate

An edge should have:

  • At least 10 joint answers
  • Recurrence on at least three collection dates
  • Coverage across at least two answer engines
  • Presence in at least three prompt clusters

Context gate

The edge should also have:

  • Manual review of at least 10 supporting passages, or all passages if fewer exist
  • One relationship label supported by at least 60% of reviewed examples
  • No single prompt cluster contributing more than half of joint support
  • Documented exceptions and secondary roles

These are operating thresholds, not scientific constants. A reputational or high-budget decision should use a larger sample, narrower scope, and stricter review.

If several analysts label relationships, double-code a sample and report agreement. Resolve disagreements by refining the label definitions rather than silently averaging conflicting judgments.

Worked example: turning 360 answers into a map

This example uses a synthetic dataset of 360 answers: 30 prompts, four answer engines, and three weekly collections. It demonstrates reproducible calculations and does not represent customer performance or an industry benchmark.

The focal product, Acme, appeared in 126 answers:

Brand paired with Acme Brand mentions Joint answers Jaccard Lift Dominant context Initial role
Rivet 114 72 0.429 1.80 54 of 72 were comparisons or alternatives Substitute
Flowbit 93 45 0.259 1.38 30 of 45 described integrations or workflows Complement
Atlas 162 54 0.231 0.95 First-named in 39 of 54 shared answers Category leader
Nova 42 21 0.143 1.43 Joint support rose from 3 to 6 to 12 Challenger candidate

For Acme and Rivet:

  • Jaccard similarity: (72/(126+114-72)=0.429)
  • Lift: ((72/360)/[(126/360)(114/360)]=1.80)
  • (P(\text{Rivet}\mid\text{Acme})=72/126=57.1%)
  • (P(\text{Acme}\mid\text{Rivet})=72/114=63.2%)

Rivet has the clearest substitute relationship because it combines high support, strong normalized overlap, lift above 1, and explicit comparison language.

Atlas has broader independent coverage and usually appears first, but its lift is close to 1. That pattern suggests a category anchor rather than Acme’s closest substitute.

Nova illustrates why trend should remain separate from size. Its support is smaller, but the pair doubled across each collection. It belongs on a watchlist; three dates are still insufficient to declare a durable market shift.

How to build and interpret the network

Represent brands as nodes and validated relationships as edges. Size nodes by answer coverage, vary edge thickness by strength, and use color or line style for relationship type. Add filters for engine, locale, prompt cluster, audience, and date.

Four patterns deserve attention:

  • Hubs: Highly connected brands may define a category or span multiple use cases.
  • Bridges: Brands connecting two clusters may enable a workflow or be expanding into an adjacent market.
  • Dense clusters: Groups with strong internal edges often form an AI-generated shortlist.
  • Isolates: Frequently mentioned but weakly connected brands may own a narrow use case.

Direction is also important. If (P(B\mid A)) is high but (P(A\mid B)) is low, B is a common reference whenever A appears, while A occupies only a small part of B’s broader context.

Hide low-support edges by default. A network containing every detected pairing becomes a visual “hairball” that obscures the relationships analysts need to inspect.

How to turn co-mentions into marketing actions

Every important edge should lead to a testable positioning, content, PR, partnership, or monitoring decision. The goal is to improve the evidence available to answer engines, not to manufacture associations by repeating competitor names on thin pages.

Observed pattern Likely interpretation Recommended action Metric to monitor
Incorrect substitutes dominate Category or audience evidence is ambiguous Clarify category, buyer, exclusions, and differentiators on product pages Correct classification rate
A rival consistently ranks first Its supporting evidence is stronger or more retrievable Publish transparent comparison criteria, proof, and evaluation content Recommendation position
A valuable complement is absent The relationship lacks crawlable documentation Create integration, partner, and workflow documentation Complement support
A leader anchors most answers The leader controls category language Differentiate around a specific audience, workflow, or outcome Shortlist inclusion beside the leader
A challenger enters more clusters The brand is gaining relevance across use cases Review its claims, sources, and audience positioning Cluster coverage and four-week support
Descriptions are inaccurate Sources contain conflicting or outdated facts Correct first-party pages and address high-impact citation sources Description accuracy
Co-mentions rise but citations do not The brand is discussed without strong source attribution Create original research, documentation, and verifiable proof Citation share

AI answers often cite more than conventional blog posts. Review the page types AI systems actually cite, including documentation, product pages, integration pages, original research, customer evidence, and transparent comparison methodology.

Follow Google’s people-first content guidance. A page created solely to place two entity names together is unlikely to provide durable value to searchers or answer engines.

How does co-mention analysis differ from AI share of voice?

AI share of voice measures visibility volume; co-mention analysis measures the structure and meaning of the brands appearing together. Recommendation rate, position, citation share, sentiment, and factual accuracy answer separate questions.

Metric Question answered
AI share of voice How often does the brand appear within the defined answer set?
Co-mention network Which brands surround it, and in what roles?
Recommendation rate How often is the brand explicitly recommended?
Position Where does the brand appear in a shortlist?
Citation share Which sources support the answer?
Sentiment Is the brand described positively, neutrally, or negatively?
Accuracy Are its category, features, audience, and limitations correct?

A brand can gain share of voice while its competitive position worsens. For example, mentions may increase mainly in answers about “cheaper alternatives to Brand X.” Conversely, lower visibility can still be valuable if it is concentrated in high-intent recommendations for the correct audience.

Use a documented AI share-of-voice calculation alongside the network to measure both volume and market context.

Which tools can support the workflow?

The right tool must preserve raw answers, prompt metadata, citations, and history—not merely display a visibility score.

Approach Best for Requirements
Spreadsheet Small exploratory studies Consistent entity and relationship coding
SQL or Python workflow Custom scoring and larger datasets Storage, extraction, quality checks, and version control
Network software Exploring clusters and centrality A clean validated edge table
AI monitoring platform Recurring multi-engine tracking Raw-answer export, prompt controls, citations, and historical comparison

Before choosing a platform, verify that it can:

  • Export full answers and prompt-level records
  • Preserve engine, locale, date, and citations
  • Separate prompt clusters and individual variants
  • Resolve aliases without merging unrelated entities
  • Filter by relationship type and recommendation position
  • Compare identical prompts over time
  • Expose sample-size and confidence warnings

For a feature and pricing overview, compare the leading AI search and LLM monitoring tools.

Common mistakes that invalidate the analysis

Most unreliable maps fail because the dataset is aggregated before entities, intents, and answer contexts are validated.

Avoid these mistakes:

  1. Counting web pages instead of generated answers
  2. Treating near-duplicate prompt variants as independent buyer demand
  3. Mixing locales, audiences, or product categories without filters
  4. Merging parent companies, products, and similarly named entities
  5. Calling every co-mentioned brand a competitor
  6. Treating first-place recommendations and passing mentions equally
  7. Using lift or percentage growth without minimum support
  8. Removing raw answers, citations, screenshots, or collection settings
  9. Comparing periods that used different prompt sets
  10. Interpreting one collection date as a trend

Co-mention is association, not causation. An answer may reveal which citation appears beside a claim, but it cannot prove which training document, retrieval step, or ranking mechanism created the relationship.

How should teams monitor and report the network?

Use a fixed baseline prompt set for trend measurement and separate exploratory prompts for discovery. Otherwise, prompt changes will be mistaken for market changes.

A defensible monthly report includes:

  • Total prompts attempted and collection success rate
  • Focal-brand answer coverage
  • Top substitute, complement, leader, and challenger edges
  • Edge strength and confidence
  • Recommendation position by engine and prompt cluster
  • New, strengthening, weakening, and disappeared relationships
  • Citation sources supporting important descriptions
  • Accuracy or category errors
  • Sample-size warnings
  • Marketing changes made during the period

Collect the baseline weekly for most categories. Daily collection is useful during launches, incidents, major model changes, or fast-moving news cycles, but it can create noise if teams react to every output variation.

Annotate product releases, content launches, PR coverage, integrations, and website changes on the timeline. An annotation creates a testable hypothesis; it does not establish that the marketing event caused the answer change.

Use an evidence-preserving AI brand-mention audit when an edge exposes inaccurate or reputationally important descriptions.

Frequently asked questions

What counts as a brand co-mention in an AI answer?

A co-mention occurs when two resolved brand entities appear within the same generated answer. Store whether they also share a list, section, paragraph, sentence, or cited passage. Answer-level pairings reveal the broad consideration set; sentence-level pairings provide stronger evidence about the relationship.

How many answers are needed for brand co-mention analysis?

There is no universal sample size. An exploratory study can begin with at least 30 intent-balanced prompts across the answer engines relevant to the audience. Keep an edge provisional until it recurs across several dates, engines, and prompt clusters and has enough supporting passages for manual review.

Does frequent co-mention mean two brands are competitors?

No. Frequent co-mention can indicate substitution, integration, category leadership, a shared attribute, or an incidental association. Inspect the prompt intent and relationship language before assigning a role. “A or B” and “A works with B” should not receive the same classification.

Can brand co-mention analysis improve ChatGPT recommendations?

It can guide improvements but cannot guarantee a recommendation. The analysis identifies relevant prompt clusters, surrounding brands, recommendation positions, descriptions, and supporting citations. Teams can then improve authoritative product, comparison, integration, documentation, and proof pages and measure whether the answer pattern changes.

How often should co-mentions be monitored?

Weekly collection is a practical baseline for most categories. Use identical prompt IDs and comparable collection conditions for trend analysis. Increase frequency around launches or fast-moving events, and require repeated observations before treating a change as durable.

Is co-mention analysis part of AI reputation management?

Yes. The brands surrounding a company influence how users interpret its category, maturity, price tier, and intended use. An incorrect recurring association can be reputationally important even when the wording is positive.

Map relationships, not just mentions

Brand co-mention analysis is useful when it explains who appears with a brand, why the pairing occurs, where each brand ranks, and whether the relationship is changing. A raw count cannot distinguish a rival from an integration partner or a category leader from a challenger.

Start with a controlled prompt set, preserve answer-level evidence, and resolve entities before aggregation. Calculate support, directional probability, Jaccard similarity, lift, placement, and recurrence. Then apply the 4R framework—Recurrence, Relevance, Relationship, and Rate of change—before treating an edge as strategically meaningful.

The final map should produce a small number of defensible actions: correct a category error, strengthen proof for an important use case, document an integration, respond to a challenger, or close a citation gap.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →