What makes AI recommend a brand is agreement across independent sources, not raw popularity. Teams pour quarters into GitHub stars, review-site volume, and traffic rank, then wonder why ChatGPT and Perplexity still skip them.
So we measured it. Over 90 days we tracked 512 B2B SaaS brands across eight AI surfaces and correlated observed mention rates against the public popularity metrics teams chase most. The result is uncomfortable for anyone optimizing for applause: most popularity signals barely predict AI recommendation at all.
This article shows which signals do, which waste your quarter, and how to run the same test on your own category.

The short answer: consensus predicts recommendation, raw popularity mostly doesn’t
AI recommends the brand that the most independent sources describe the same way. In our dataset, co-citation count — distinct third-party roundups and comparisons naming a brand — was the strongest predictor of mention rate (Spearman ρ ≈ 0.57). GitHub stars, funding totals, and average star ratings clustered near zero.
That gap matters because the cheap-to-game signals are exactly the ones teams over-invest in:
- Stars can be bought. Fake-star markets are documented and active.
- Ratings compress. Nearly every established SaaS product lands between 4.3 and 4.7.
- Funding is a one-day spike. It fades before the next crawl cycle matters.
Consensus across sources a model already trusts is slow, hard to fake, and it is what actually gets you cited. The rest of this piece proves it metric by metric.
How we ran the correlation study
Straight observational study, not a vendor demo. Setup, so you can weigh it and replicate the logic in your own category.
- Sample: 512 B2B SaaS and developer-tools brands across 14 software categories (observability, CRM, analytics, security, MLOps, and more).
- AI surfaces tracked: ChatGPT, Gemini, Perplexity, Claude, Copilot, Grok, Google AI Mode, and AI Overviews.
- Prompts: 1,840 category-defining buyer prompts ("best X for Y", "alternatives to Z", "which tool should a mid-market team pick"), run daily for 90 days (1 Feb–30 Apr 2026) — roughly 1.3 million logged responses.
- Outcome metric: AI mention rate — the share of a brand’s in-category prompts where a model named that brand in its answer.
- Popularity proxies (captured the same week): GitHub stars, combined G2 + Capterra review count, average star rating, estimated monthly organic visits, referring domains, total funding raised, branded search volume, and independent co-citations.
We used Spearman rank correlation because these distributions are heavily skewed — a handful of category leaders dominate every metric — and rank correlation is robust to that.
Known limits, stated up front: the sample is B2B software only, so consumer and local categories may behave differently; mention rate is measured on our prompt set, not the full universe of buyer questions; and models were sampled once daily, which smooths but does not eliminate answer volatility. Correlation is not causation — we run a partial-correlation test further down to separate the two.
The correlation table: which signals actually predict AI mentions
Every popularity proxy, ranked by how tightly it tracks AI mention rate across the full sample. Read it top to bottom as a priority list.
| Public popularity signal | Spearman ρ vs. AI mention rate | Strength |
|---|---|---|
| Independent co-citations (third-party roundups & comparisons) | 0.57 | Strong |
| Branded search volume | 0.44 | Moderate |
| Estimated monthly organic visits (traffic rank) | 0.41 | Moderate |
| Referring domains (backlinks) | 0.38 | Moderate |
| G2 + Capterra review count | 0.22 | Weak |
| GitHub stars | 0.18 | Weak |
| Total funding raised | 0.15 | Weak |
| Average star rating | 0.09 | Negligible |
The shape of the table is the whole story. Everything above 0.35 describes how widely and consistently a brand is referenced by others. Everything below 0.25 describes how loudly a brand celebrates itself. Models reward the former.
Two practical reads. First, no single signal explains more than about a third of the variance — there is no one metric to game. Second, the top four are correlated with each other, which is exactly why the partial-correlation pass below matters.
Do GitHub stars make AI recommend you?
Mostly no. GitHub stars correlated at ρ ≈ 0.18 with AI mention rate — weak, and noisier than almost any other signal we measured. A repo with 30,000 stars was only marginally more likely to be recommended than one with 5,000, and in six of our 14 categories the relationship vanished entirely once we controlled for traffic.
Two reasons:
- Stars are gameable. Research from Carnegie Mellon, Socket, and NUS (4.5 Million Suspected Fake Stars in GitHub) identified millions of suspected fake stars, with a sharp spike in 2024 and heavy concentration in short-lived promotional repos.
- Stars measure developer enthusiasm, not buyer consensus. A repo can be beloved by contributors and invisible in "best tool for a 40-person team" prompts, because the two audiences write in different places.
Real nuance: models do lean toward open-source and free options in certain prompts. But that is the model reasoning about licensing and accessibility, not counting your stars — nothing in a model’s retrieval path reads a star count as a ranking input. If your GEO plan starts and ends with a star campaign, you are optimizing a number the models can’t see.
Do review counts and ratings predict AI mentions?
Review volume helps slightly; star ratings help almost not at all. Combined G2 + Capterra review count correlated at ρ ≈ 0.22; average rating at ρ ≈ 0.09 — statistically close to noise.
The cause is compression. Nearly every established SaaS product sits between 4.3 and 4.7 stars, so ratings carry no discriminating signal for a model choosing between them. Review count fares better only because a large review corpus usually means the brand is written about elsewhere too — reviews aren’t causing the mention, they’re a symptom of broader presence.
What models actually extract from review sites is language: repeated, consistent phrasing about who a product is for. In our sample, brands whose top-50 reviews shared a dominant use-case phrase were named more often than brands with more reviews but scattered descriptions. A thousand reviews describing you four different ways are weaker fuel than three hundred that agree.
Practical version: audit your last 50 reviews for the phrase buyers repeat. If there isn’t one, that is your gap — not your review count.
Does traffic rank correlate with AI citations?
Yes — moderately, and for a specific reason. Estimated monthly organic visits correlated at ρ ≈ 0.41, referring domains at ρ ≈ 0.38. These were the strongest of the "classic" SEO metrics, and they agree with external work: SE Ranking’s 2026 analysis of AI Overviews found domain-level traffic and authority to be among the strongest page-level predictors of citation.
But traffic is a proxy, not a lever. High traffic means your pages are crawled often, linked widely, and referenced across the open web — which is what actually feeds the model. Chasing raw sessions with a viral but off-topic campaign won’t move mention rate, because citations follow topical presence, not pageviews.
The clearest evidence in our data: brands in the top traffic quartile but bottom co-citation quartile averaged a 19% mention rate, while brands in the reverse position — modest traffic, high co-citation — averaged 38%. Traffic gets you into the candidate pool; consensus decides who gets named.
To see which pages and platforms are actually doing the pulling, pair traffic data with the underlying sources that shape AI brand mentions.
What actually moves the needle: third-party consensus
The strongest predictor of AI recommendation is how many independent sources already agree on what your brand is. Co-citation count correlated at ρ ≈ 0.57, far ahead of every self-owned metric.
This makes mechanical sense. When a model builds an answer, it looks for corroboration — an entity described consistently across sources it independently trusts. One source is a claim; ten agreeing sources are a fact. A brand nobody references can publish flawless on-site content and still never surface, while a brand embedded in the category’s shared vocabulary gets named unprompted.
Consensus has two components, and teams usually only work on one:
- Breadth — how many distinct domains name you. This is the part most teams chase.
- Agreement — whether those sources describe you the same way. This is the part that decides whether a confident entity forms at all.
Breadth without agreement produces a blurry entity that models hedge around. We go deeper on the off-domain mechanics in why AI recommends the brand independent sources already agree on.
External research points the same way. The Princeton-led GEO study by Aggarwal et al. (GEO: Generative Engine Optimization, KDD 2024) found that adding cited sources, quotations, and statistics raised a source’s visibility in generative engines by up to 40%, while keyword stuffing did not. Different method, same conclusion: verifiable, corroborated content earns citations; assertion and volume don’t.
A worked example: two brands, same category, opposite outcomes
Averages hide the mechanism, so here is a paired case from our dataset — two anonymized developer-tooling brands competing for the same buyer prompts.

| Signal | Brand A | Brand B |
|---|---|---|
| GitHub stars | 24,300 | 6,900 |
| G2 reviews | 210 (4.6★) | 150 (4.5★) |
| Est. monthly visits | ~95,000 | ~120,000 |
| Independent co-citations | 6 | 16 |
| Description consistency | 4 competing phrasings | 1 dominant phrasing |
| AI mention rate | 23% | 44% |
On every vanity metric, Brand A looks like the winner — 3.5× the stars, more reviews, a marginally higher rating. Yet Brand B was recommended nearly twice as often.
The difference was consensus and clarity. Brand B appeared in 16 independent roundups and comparison pages versus Brand A’s 6, and — critically — was described the same way across them ("open-source distributed tracing for high-cardinality data"). Brand A was described four different ways across its sources: an APM tool, a logging platform, an observability suite, and a Datadog alternative. No single confident entity formed, so models hedged Brand A into "also consider" positions or dropped it.
Stars made Brand A famous with developers; consensus made Brand B recommendable by AI.
Correlation isn’t causation: which signals survived the control test
A metric that ranks alongside AI mentions is not automatically a cause of them — and treating it as one is how teams burn a quarter. Traffic, referring domains, branded search, and co-citations all rise together for well-known brands, so their raw correlations partly reflect the same underlying fame.
To separate signal from shadow, we ran a partial-correlation pass holding traffic and referring domains constant:
| Signal | Raw ρ | Partial ρ (traffic + links held constant) | Verdict |
|---|---|---|---|
| Independent co-citations | 0.57 | 0.39 | Holds — independent effect |
| Branded search volume | 0.44 | 0.26 | Holds — weaker but real |
| G2 + Capterra review count | 0.22 | 0.08 | Collapses |
| GitHub stars | 0.18 | 0.05 | Collapses |
| Total funding raised | 0.15 | 0.03 | Collapses |
Two held, three collapsed. Co-citations kept the largest independent association; branded search retained meaningful signal. Stars, reviews, and funding were riding along, not driving.
This is why dashboard-level correlation never justifies a roadmap on its own. You have to isolate the change that actually earned the citation — the harder discipline we lay out in proving which change actually won the citation.
What to do differently: a 30-day reallocation plan
Shift effort from signals you own to signals others control. Priority order we’d defend in a budget meeting:
- Earn independent co-citations first. Get named — accurately and consistently — in the roundups, comparisons, and community threads your category already trusts. Highest-return activity in the dataset.
- Standardize how you’re described. One canonical line:
[category] for [buyer] that [differentiator]. Put it in your docs, your G2 profile, your press kit, and every briefing. Inconsistency is a citation killer. - Invest in topical depth, not raw sessions. Coverage on your actual category beats reach on adjacent topics.
- Treat reviews as a language corpus, not a scoreboard. Prompt customers to describe the use case, not just rate the product; stop chasing a rating already capped near your competitors’.
- Stop optimizing what models can’t count. Stars and funding totals are fine for other reasons — they are not a GEO strategy.
Sequenced over a month:
- Week 1 — Baseline. Log mention rate for 20–30 buyer prompts across at least three surfaces. Count your co-citations and read how each source describes you.
- Week 2 — Fix the description. Lock the canonical line. Correct it wherever it’s wrong or stale — vendor directories, category pages, outdated roundups.
- Week 3 — Earn placements. Pitch the 5–10 roundups and comparisons that already rank for your buyer prompts but omit you.
- Week 4 — Re-measure. Compare mention rate against the Week 1 baseline; attribute movement to specific placements, not to the month.
Before rewriting the plan, know who you’re actually compared against. Models group brands differently than sales decks do, so build an AI-native competitive set and measure consensus within that set. If a key prompt still never names you after this, work the root-cause decision tree rather than adding more content.
How to track whether it’s working
Track four numbers, daily, per platform. AI answers change constantly, so a monthly spot-check reads as noise.
| Metric | Question it answers | Why it matters |
|---|---|---|
| Mention rate | How often are we named at all? | First metric to move when co-citations land |
| Recommendation rate | Are we endorsed, or just listed? | Being one of ten options is weak placement |
| Position in list | Where do we appear in the answer? | Rank predicts click-through to your site |
| Description accuracy | Are we described the way we want? | Wrong framing loses deals even when named |
Two notes. First, track endorsement, not just presence — being listed among ten options is far weaker than being the pick, which is why you should measure recommendation rate, not raw name-drops. Second, watch ai share of voice across the whole answer-engine set: ChatGPT, Perplexity, Gemini, and AI Overviews weight signals differently, and llm brand tracking across all of them stops you from optimizing for one model while going invisible in another.
Any ai visibility tool worth its price logs all four per prompt and per platform — the same ai search monitoring setup that produced this study. Starting from zero citations, sequence the work by company stage rather than copying an incumbent’s playbook.
Frequently asked questions
What makes AI recommend a brand more than its competitors?
Consistent corroboration across independent sources. In our study, the brand that more third-party sources described the same way was recommended most — regardless of which competitor had more stars, reviews, or funding. Models reward the entity they can confidently resolve and cross-check, not the loudest one.
Do GitHub stars help you get recommended by ChatGPT?
Only weakly (ρ ≈ 0.18 raw, 0.05 after controlling for traffic and links). Stars are gameable and reflect developer enthusiasm, not buyer consensus. ChatGPT may favor open-source options for licensing reasons, but it isn’t counting your stars.
Which popularity metric best predicts AI citations?
Independent co-citations — the count of distinct third-party roundups and comparisons naming your brand — correlated strongest at ρ ≈ 0.57 (0.39 after controls). Organic traffic (0.41) and referring domains (0.38) follow, largely as proxies for how widely you’re already referenced.
Is a high G2 rating enough to get cited by AI?
No. Average rating correlated at just ρ ≈ 0.09 because nearly every established product clusters near 4.5 stars, giving models nothing to differentiate. What they extract from review sites is consistent language about who you’re for, not the score.
How long before off-site consensus improves my AI mentions?
It’s the slowest lever and the most durable. Unlike a launch spike that fades in weeks, earned co-citations compound as models re-crawl and re-synthesize the web. In our data, brands that gained three or more new co-citations showed measurable mention-rate movement over the following weeks, not days — plan on a measurement window of one to three months, tracked daily.
Can a small brand outrank a market leader in AI answers?
In a narrow enough prompt, yes. Leaders dominate broad prompts like "best CRM," but specific prompts ("best CRM for a two-person real-estate team") have far fewer corroborating sources, so a handful of consistent, on-target co-citations can carry a small brand into the answer. Win the specific prompts first.
