AI recommendation bias is the systematic tilt that decides which brands an AI engine names — before you optimize a single page. Over 90 days we tracked 214 brands across eight AI engines, and the pattern was hard to miss: most shortlists are settled by four structural signals — company size, price model, source concentration, and news recency — long before content quality enters the picture.
This is a field study, not a hot take. Below are the effect sizes we measured per engine, how they line up with independent academic research, and the playbook we use to move a brand up the shortlist despite the tilt.

What is AI recommendation bias?
AI recommendation bias is the tendency of a generative engine — ChatGPT, Gemini, Perplexity, Claude, Copilot, Grok, or Google’s AI surfaces — to name some brands more often than their merit justifies, driven by structural signals rather than answer quality. It is a commercial-visibility tilt in the output, not the demographic-fairness sense of the word. It shows up as a lopsided AI share of voice: a handful of names fill the answer, and everyone else stays invisible.
The reason it matters more than classic search bias is simple math. A results page shows ten blue links; an AI answer names three to five brands and stops. The tilt decides those slots, and there is no page two.
How we ran the field study
We measured unprompted inclusion — how often a brand appears in an AI answer when the user never types its name. That is the honest test of whether an engine already carries a bias toward you.
- Brands: 214, spanning B2B SaaS, developer tools, and consumer tech
- Queries: 26 buyer-intent prompts — "best X for Y," "top tools for Z," "alternatives to…"
- Engines: 8 — ChatGPT, Gemini, Perplexity, Claude, Copilot, Grok, Google AI Overviews, and Google AI Mode
- Window: Q2 2026 (April–June), sampled multiple times daily
- Volume: ~590,000 answer snapshots
- Metric: Unprompted Inclusion Rate (UIR), plus a tilt multiplier — the UIR ratio between the advantaged and disadvantaged group on each axis
These numbers come from our own tracking, not a controlled lab, so treat them as directional. We publish the method so you can reproduce the shape of the finding against your own brand, not just take our word for it.
The four tilts that decide the shortlist
Four signals explained most of the variance in who got named — and, critically, they operate before any answer engine optimization work begins. Here is the map, then the axis-by-axis detail.
| Tilt | What it rewards | Strongest on | Weakest on |
|---|---|---|---|
| Company size | Established, evidence-rich incumbents | ChatGPT, AI Overviews | Perplexity |
| Price model | Free & open-source options | Perplexity, Claude | AI Overviews, ChatGPT |
| Source concentration | Brands named by a few dominant domains | Perplexity, AI Overviews | Claude |
| News recency | Recent funding or launch coverage | Perplexity, AI Mode | Claude |
No single engine tilts on everything. That is the practical opening — different engines reward different signals, so the fix is never one-size-fits-all.
Tilt 1 — Company size (the incumbent tilt)
Large, established brands were named 2.0×–3.1× more often than equally relevant challengers, and the gap was widest on ChatGPT. This is the tilt most people mean when they ask whether AI favors big brands.
| Engine | Incumbent tilt multiplier |
|---|---|
| ChatGPT | 3.1× |
| Google AI Overviews | 2.9× |
| Gemini | 2.6× |
| Copilot | 2.5× |
| Google AI Mode | 2.3× |
| Claude | 2.2× |
| Grok | 2.0× |
| Perplexity | 1.5× |
Independent research is even starker at the extreme. In a 2026 analysis of LLM recommendation systems, when competing products had identical specifications, models picked the recognized brand in 100% of 670 trials — yet the same study found brand identity alone explained just 1.2% of ranking variance, while product parameters explained 82.4%, per the arXiv "Incumbent Advantage" study on brand bias in LLM recommendations. A rating edge as small as +0.075 stars — less than the gap between a 4.3 and a 4.4 — was enough to flip the pick.
The takeaway reframes the whole debate. The incumbent tilt is a proxy for evidence density, not a hard love of bigness — big brands simply carry more citations, reviews, and comparisons for the model to lean on. We unpack the mechanism in our breakdown of whether ChatGPT favors big brands.
Tilt 2 — Price model (the free-and-open-source tilt)
When a prompt did not mention budget, free and open-source tools filled 37% of shortlist slots on average, peaking at 48% on Perplexity. Paid products were quietly penalized for a signal they never chose.
The mechanism is coverage, not ideology. Open-source tools accumulate dense documentation, GitHub activity, and forum threads — exactly the corpus an engine reaches for when it wants a "safe," well-attested answer. Free tiers also read as low-risk recommendations, which the model treats as a feature.
Slot share for free/OSS options ran highest on Perplexity (48%) and Claude (44%) and lowest on Google’s AI surfaces and ChatGPT (29–33%). If you sell a paid product, this tilt is beatable — but only deliberately, by making your paid value legible and your risk reversible.
Tilt 3 — Source concentration (the citation-cluster tilt)
In most categories, the top three domains supplied more than half of all citations — 61% on Perplexity and 58% on Google AI Overviews. A small cluster of pages decides who is quotable.

The usual suspects were Reddit threads, Wikipedia, one or two review platforms like G2, and a couple of "best of" listicles. If your brand is absent from that cluster, you are structurally hard to cite — no amount of on-site copy compensates. This is where AI citations are won or lost, and it is why our study of the pages cited by every engine matters more than domain authority.
Concentration is a retrieval artifact: embeddings, chunking, and reranking all reward dense, well-structured passages, so the same few sources keep surfacing. Get into that cluster or stay invisible.
Tilt 4 — News recency (the recency tilt)
A funding round or product launch lifted unprompted inclusion by up to 14 percentage points — but the lift decayed fast, with a half-life of about 9 days on live-retrieval engines.
That decay curve is the part most PR teams miss. On Perplexity and Google AI Mode, a launch spike faded to a few points within two weeks. On ChatGPT, the same event moved the needle less at peak (+6 pp) but stuck around far longer — a half-life closer to 34 days — because it leans on slower-moving training and memory rather than the live index.

The operational lesson: refresh coverage before the half-life expires, and don’t judge a launch by day-one numbers alone.
Effect sizes per engine: which engine tilts hardest?
No engine is neutral, but each tilts on a different axis. Perplexity leans on recency and concentrated sources because it retrieves live and cites as it answers; ChatGPT and AI Overviews lean on size and incumbency; Claude is the most merit-stable but cites the least, so it is hardest to influence with fresh content.
| Engine | Company size | Price model | Source concentration | News recency |
|---|---|---|---|---|
| ChatGPT | High | Low | Medium | Low–Med |
| Gemini | High | Medium | Medium | Medium |
| Perplexity | Low | High | High | High |
| Claude | Medium | High | Low | Low |
| Copilot | Med–High | Low | High | Medium |
| Grok | Medium | High | Medium | High |
| Google AI Overviews | High | Low | High | Medium |
| Google AI Mode | Med–High | Low | High | High |
There is a second-order tilt hiding inside every shortlist: order. Columbia Business School researchers found that across 5,447 prompts, AI systems chose the first option listed 63% of the time, regardless of wording, per Columbia’s research on ChatGPT’s bias for the first option. So being named is only half the battle — slot #1 compounds the tilt. Engines also disagree on who to name, which we quantify in how much ChatGPT, Perplexity, and Gemini overlap on brand picks.
Why the tilts exist — evidence density, not favoritism
The engines are not playing favorites; they are pattern-matching on signals that correlate with a safe, verifiable answer. Every tilt is a shortcut to "this brand is real, stable, and won’t embarrass me."
That is why supplying the evidence directly works. Landmark GEO research found that adding citations, quotations, and statistics boosted a source’s visibility in generative answers by up to 40% versus generic SEO-style optimization, per the GEO study by Aggarwal et al.. Read the four tilts as proxies and the strategy writes itself:
- Size proxies for stability → publish proof of scale and named customers.
- Free proxies for low risk → make your paid value legible and risk-reversible.
- Concentrated citations proxy for consensus → get into the cluster the engine already trusts.
- Recency proxies for relevance → keep a steady cadence of citable news.
Beat the bias by feeding the underlying evidence, not by fighting the signal.
How to counter AI recommendation bias: a playbook
You cannot delete the tilt, but you can supply the signals it rewards. Here is the sequence we run, in order of use.
- Baseline first. Measure your Unprompted Inclusion Rate and AI share of voice per engine with an AI visibility tool — you cannot fix a tilt you have not sized.
- Enter the citation cluster. Earn placement on the two or three domains that supply most of your category’s citations — listicles, review platforms, and eligible Wikipedia entries.
- Publish evidence-dense pages. Add statistics, named sources, and honest comparisons — the core of both answer engine optimization and generative engine optimization. Start with our practical definition of answer engine optimization.
- Time content to news windows. Ship citable updates and refresh them before the recency half-life decays the lift.
- Build entity and author authority. Named experts and a clear company entity act as a proxy for the size signal you may not have yet — see how named experts and bylines earn AI citations.
- Own the objection turn. Address downsides and "who is this not for" directly, so the engine can quote your honest answer instead of a competitor’s — the tactic we detail in winning the objection turn in AI chats.
- Monitor continuously. Watch brand mentions across ChatGPT and every other engine on a regular cadence, and treat AI visibility as ongoing maintenance, not a one-time audit.
The brands that get recommended by ChatGPT are rarely the biggest in the room. They are the ones that hand the engine the exact evidence its tilt is hunting for — and keep the tracking on to prove it worked.
Frequently asked questions
Is AI recommendation bias the same as bias in the training data?
No. Training-data bias is baked into the model’s weights, while AI recommendation bias is what you observe in the output — which brands get named for a query. Our field study measures the output tilt, because that is what marketers can actually move with content, citations, and timing.
Which AI engine is least biased toward big brands?
In our data, Perplexity showed the smallest incumbent tilt (1.5×) because it retrieves live and cites, letting a well-structured challenger page break in. Claude was the most merit-stable overall but the hardest to influence, since it cites fresh sources the least.
Can a small brand overcome AI recommendation bias?
Yes. Independent research shows a rating edge of roughly +0.075 stars — or clearly superior specs — can flip an LLM’s pick away from a known incumbent. The lever is evidence density (citations, statistics, and reviews), not raw company size.
How do I measure AI recommendation bias for my own brand?
Track your Unprompted Inclusion Rate across every engine over time, ideally daily. A dedicated AI visibility tool compares your AI share of voice against competitors and flags which of the four tilts is holding you back on each platform.
Does rewording the prompt remove the bias?
Only partly. Columbia’s research found that aggregating many differently worded prompts cancels the order bias, but structural tilts — size, price, sources, recency — persist across phrasings. The durable fix is changing the signals, not the wording.
