AI Wrong Company Size Fit: How Answer Engines Decide Who Your Product Is For

by

·

AI Wrong Company Size Fit: How Answer Engines Decide Who Your Product Is For

AI wrong company size fit happens when ChatGPT, Gemini, Perplexity or AI Overviews assigns your product to a buyer tier you don’t serve — or to no tier at all. Your features get described correctly. The audience label is wrong, and that label decides which shortlists you’re eligible for.

The failure is silent. You never show up in "best tool for a 20-person startup," you never show up in "enterprise-grade options," and nobody sees a ranking drop — because there was no ranking to drop.

We ran a 28-day panel on this: 84 size-tier prompts, five engines, 11,760 answers. Two findings drove everything below. Tier phrasing changes the shortlist more than category phrasing does — startup-tier and enterprise-tier answers overlapped by only 19%. And the most common failure isn’t being put in the wrong tier; it’s being put in none.

Side-by-side AI answers for the same software category showing completely different brand shortlists for a startup-tier prompt and an enterprise-tier prompt, illustrating AI wrong company size fit

What is size-tier misclassification?

Size-tier misclassification is when an answer engine infers the wrong customer size for your product and files you under a segment you don’t sell to. It is not a factual hallucination — your feature list is fine. The error sits in the audience label.

Tier works as a filter, not a tiebreaker. When a buyer types "for a team of 15," the model does not rank every product in the category and pick the best five. It first narrows to products it believes serve teams of 15, then ranks inside that set. Wrong label, and you’re excluded before quality is ever considered.

Three consequences:

  • Feature depth doesn’t rescue you. You’re not in the candidate pool being compared.
  • Review count doesn’t rescue you. Volume is a within-set signal.
  • Your Google rankings tell you nothing about it. The query that excluded you may never have been typed into a search box.

Why answer engines assign a size tier at all

Engines assign a tier because buyers ask for one. Roughly half the commercial prompts we track in B2B software categories carry a size, budget, or team-shape qualifier — "small team," "solo," "mid-market," "enterprise-ready," "for a 200-person company," "cheap." Those qualifiers survive into retrieval, so the model needs passages that match them.

The mechanism is ordinary retrieval, not judgment. The model matches the qualifier against text chunks it can find, then writes a summary that reads like an opinion. If your site never states a team size in the same passage as your product name, there is nothing to match — so the engine borrows a size from adjacent evidence: your pricing page, your logo strip, or a directory’s filter label.

Chunking is why co-location matters. Retrieval systems split pages into passages and embed each one separately, so a size claim in your footer and a product claim in your hero may never travel together in the same retrieved chunk. Presence somewhere on the domain is not the same as presence in one retrievable passage.

What we found: 11,760 answers on size-tier prompts

Methodology first, so you can weigh the numbers.

What we ran. MaxAEO tracked 84 prompt strings — 12 phrasings each across 7 B2B software categories (project management, CRM, HR/payroll, analytics, help desk, security posture, data integration). Each category set covered three tiers: startup/small team, mid-market, enterprise. Every prompt ran on five engines (ChatGPT, Gemini, Perplexity, Claude, Google AI Mode) once daily for 28 days, 12 May to 8 June 2026. That’s 11,760 answers.

Who we watched. 42 named B2B SaaS products that publicly position as mid-market — they either state a range spanning SMB and enterprise, or publish both a self-serve tier and a "contact sales" tier.

What we measured. For each brand, a tier-correct mention rate: the share of prompts in that brand’s own stated tier where the brand was named at all. We logged cited sources, and re-asked each engine to justify its pick ("why is this a fit for a team of 20?") once per week.

Tier phrasing splits the shortlist almost completely

The same category question, asked with a different size qualifier, returned a substantially different set of brands.

Prompt tier Median brands named per answer Overlap with the opposite-tier shortlist
"for a 5–20 person startup" 6.2 19%
"for a ~200-person company" 5.8 34%
"for a 5,000-employee enterprise" 5.1 19%

Only 19% of brands named in startup-tier answers also appeared in enterprise-tier answers for the same category. Engines treat the two ends as near-disjoint sets. The mid-tier prompt sat in between — it borrowed from both, but it borrowed the already-named brands, not the unnamed ones.

The mid-market void: excluded from both, not misplaced in one

The finding we didn’t expect. Of the 42 mid-market brands, 26 (62%) appeared in fewer than 10% of answers at both ends — not pushed up into enterprise, not pushed down into SMB, simply absent from tier-qualified answers while still appearing in generic "best X tools" answers for the same category.

That’s the trap. Teams assume misclassification means "AI thinks we’re enterprise when we’re not." More often it means AI has no size hypothesis about you at all, leaving you eligible only for the shrinking pool of unqualified questions.

The worst performers were explicit about it. Five brands used "any size," "from solo founders to the Fortune 500," or "teams of all sizes" as their primary audience statement. Their median tier-correct rate was 7% — the lowest cohort in the study. Claiming everyone reads, to a retrieval system, as claiming no one.

Which evidence engines cited to justify size fit

When we asked engines to explain a size recommendation:

Evidence cited in size-fit justifications Share of justifications
Public pricing page (tier names, seat minimums, free tier) 68%
Dedicated segment or solutions page 47%
Third-party review-site or directory tier labels 41%
Customer stories naming a company of comparable size 34%
Docs, onboarding time, or implementation claims 23%
Homepage hero copy alone 11%

One caveat we hold to: a model’s stated reasons are not ground truth about its retrieval. Post-hoc justifications are directional evidence about what the model finds legible, not a log of what it fetched. Read these as a ranking of signal legibility. Causation we tested separately, below.

Three correlations from the same dataset:

  • Brands with no public pricing were named in enterprise-tier answers 2.4× more often than in startup-tier answers.
  • Brands with a public free tier flipped it — 3.6× more startup-tier mentions than enterprise-tier ones.
  • Brands showing four or more recognizable enterprise logos above the fold had "enterprise" in 71% of their generated brand descriptions, versus 22% for brands with no logo strip.

Your pricing page is doing audience segmentation whether you meant it to or not.

The three misclassification modes

Diagnose which one you have before writing a single new page. The fixes differ, and applying the wrong one makes it worse.

Mode What the AI says Usual cause First fix
Overshoot — read as enterprise-only "Best for large organizations with dedicated admin resources" Hidden pricing, enterprise logo strip, SOC 2 / SSO messaging dominating the homepage Publish a seat-banded entry tier and one small-team customer story
Undershoot — read as an SMB toy "Good for small teams, but larger organizations may outgrow it" Free tier headlining, no security/compliance page, no named large customers Ship a scale-evidence page: limits, admin controls, migration path, one 500+ seat example
Void — no tier assigned, excluded from both Absent from tier-qualified answers, present in generic ones "Teams of all sizes" language, one undifferentiated product page, no numbers anywhere Split into per-tier segment pages with explicit employee ranges

Most mid-market products in our panel were in Void, not Overshoot. Check before assuming.

Diagnostic table mapping the three AI size-tier misclassification modes — overshoot, undershoot and void — to their causes and first fixes

The Size Signal Stack: five layers AI reads to size you

Ordered by how often changing them moved a tracked mention rate in our follow-up test — not by effort.

1. Price shape. Strongest signal by a wide margin. Not the number — the shape: is there a public tier, is there a seat minimum, does the cheapest plan floor at 5 seats or 50? A "contact sales" wall reads as enterprise even when your median deal is 30 seats.

2. Named proof at the right scale. One customer story saying "a 40-person agency" does more tier work than ten saying "a leading provider." Engines can only match sizes they can read. Put the employee count in the same sentence as the customer name and the outcome.

3. Segment pages with numbers in them. One page per tier you actually serve, each stating the range in its opening lines. Not "for growing teams" — "for companies between 50 and 500 employees."

4. Third-party tier labels. Review-site categories, directory filters, and competitor comparison pages carry real weight — 41% of justifications leaned on them. This is the layer you don’t control directly, and it’s why a competitor’s "alternatives" page can end up as AI’s main source about you. Fix your own profiles before writing new content.

5. Machine-readable audience. Underused and cheap. Schema.org defines BusinessAudience for describing the organizations a product serves, including numberOfEmployees as a min/max range, and Google documents structured data for software applications that it attaches to. Markup won’t override contradicting body copy — but where copy is ambiguous, it’s the cheapest disambiguation you can ship.

Layer 5 without layers 1–3 does nothing. Layers 1–3 without layer 4 stall, as the test below showed.

How to fix AI wrong company size fit in 30 days

Each step produces evidence the next one depends on.

  1. Baseline the tier prompts, not the brand prompts. Write 12 prompts per category: four phrasings each for the tier below you, your tier, and the tier above. Run them across every engine your buyers use and record who gets named. This is the number you’ll defend later.
  2. Classify your mode. Overshoot, Undershoot, or Void, using the table above. If you’re absent from all three tiers and from the generic question, you have a category problem, not a size problem — solve that first.
  3. State the range in your first 100 words. Homepage and every segment page. Use an employee or seat number. "Built for teams of 50 to 2,000" beats every adjective available to you.
  4. Fix the price shape. Publish at least one seat-banded tier matching your true entry point. If you can’t publish numbers, publish the band — "starts at 25 seats" is a retrievable fact; "contact us" isn’t.
  5. Write one customer story per tier, with company size in the headline. Two stories beat one whitepaper here.
  6. Update third-party profiles. Review sites, directories, and your own comparison pages. Make the size filters you’re listed under match the pages you just shipped.
  7. Add audience markup with a numberOfEmployees range on product and pricing pages.
  8. Re-run the baseline weekly for six weeks. Tier movement is slower than brand-mention movement — most test brands saw nothing for 12–18 days, then a step change.

Ordering matters more than completeness. Shipping schema in week one and pricing changes in week five inverts the use.

What happened when nine brands shipped the full evidence set

Nine panel brands implemented steps 3 through 7 and let us keep tracking. After six weeks:

  • Median tier-correct mention rate: 12% → 34%.
  • Seven of nine improved, median 12% → 39%.
  • Two of nine were flat (11% → 12%).

The two failures are the useful part. Both had done the on-site work well — segment pages, seat bands, schema, customer stories. Both were still described as enterprise-only across three engines. The difference: neither had updated their third-party profiles, where they stayed categorized under enterprise-only filters, and both had a competitor comparison page ranking well that described them as "aimed at large deployments."

Their own site said one thing. The corroborating web said another. The engines went with the corroboration. Self-description sets the hypothesis; third-party agreement confirms it; confirmation is what survives into the answer. If you want to attribute a specific tier change to a specific edit rather than to ambient drift, the same discipline used for proving causality in AI search applies — change one signal at a time and hold a control prompt set.

Limitations: nine brands is a small sample, a six-week window can’t separate our changes from ambient model updates, and we ran no matched control group. Treat 12% → 34% as an observed direction with a plausible mechanism, not a guaranteed multiplier.

Does this apply outside B2B SaaS?

Our panel was B2B software, so the pricing-page finding is strongest there. The mechanism generalizes anywhere buyers qualify by scale, and the signals shift:

  • Agencies and services. Tier is read from case studies and client logos, not pricing. Publish engagement size ("$15k–$60k projects," "teams of 3 to 12") in the same passage as the service name.
  • Hardware and infrastructure. Capacity numbers do the tier work — seats, nodes, throughput ceilings. A stated maximum reads as an upper bound on your audience; state it deliberately.
  • Regulated categories. Compliance pages act as tier signals in both directions. SOC 2 and HIPAA copy pushes toward enterprise; a free tier with no DPA pushes toward SMB.

One thing does not generalize: the 12% → 34% number. That was nine B2B SaaS brands over six weeks. Re-baseline in your own category before assuming the magnitude transfers.

How to know the fix actually worked

Tier mention rate is the metric — the share of your tier’s prompts where you’re named. Not traffic, not overall share of voice, not sentiment.

Track it three ways:

  • Inclusion: named at all in tier-qualified answers.
  • Position: where in the list — first-three placement behaves very differently from a trailing mention.
  • Description: the adjective attached to you. "Enterprise-grade" showing up in answers to startup prompts means the tier label moved but the framing didn’t.

One more control: check who else gets named alongside you. If your tier-correct rate rises but the co-mentioned brands are all a tier above, the engine has moved you into the wrong neighborhood. Building an AI-native competitive set from what ChatGPT actually names gives you that read directly — your peer list is a faster tier diagnostic than your own mention rate.

Sampling frequency matters for a specific reason: tier classifications are stickier than brand facts, but they do flip, usually after a pricing change or a widely-syndicated review update. Weekly sampling catches the flip; quarterly reporting catches it a quarter late. Any of the AI search monitoring tools that track prompt-level mentions can run this, provided you can store your own tier-qualified prompt set rather than only brand-name prompts.

Mistakes that keep the wrong tier stuck

Adding a new segment page without changing the price shape. The pricing page outvoted the segment page in every case we watched. Fix the stronger signal first.

Writing "for teams of all sizes" on the new segment pages too. We saw this three times. The page exists, the range doesn’t, nothing changes.

Treating it as a copywriting problem. Adjectives are not retrievable claims. Numbers, named customers, and published bands are.

Optimizing for the tier you want instead of the one you serve. If AI moves you into enterprise answers and your product genuinely can’t handle 5,000 seats, you’ve bought unqualified demos and future churn. This work should make the label accurate — which is also what Google’s guidance on creating helpful, reliable, people-first content asks for.

Assuming one engine speaks for all of them. Tier-correct rates varied by up to 24 points across engines for the same brand in the same week. Fixing ChatGPT is not fixing AI search.

Frequently asked questions

How do I know if AI has the wrong company size fit for my product?

Ask three engines the same category question three ways: for a 10-person team, a 200-person company, and a 5,000-employee enterprise. If your brand appears in none of them but does appear in the unqualified "best tools for X" version, you have a Void-mode problem. If it appears only in the wrong one, you have Overshoot or Undershoot.

Does hiding pricing really make AI think we’re enterprise?

In our panel, yes — brands without public pricing were named in enterprise-tier answers 2.4× more often than in startup-tier ones. It’s a correlation, not proof of causation, and the brands hiding pricing skewed enterprise to begin with. The mechanism is straightforward: a "contact sales" page gives retrieval nothing to match against a seat count, so the model uses the absence itself as the signal.

How long does it take to correct a size-tier misclassification?

In the nine-brand test, most saw no movement for 12–18 days, then a step change over the following two to three weeks. Third-party profile updates propagated more slowly than on-site changes. Plan on six weeks minimum before judging the work, and keep measuring weekly so you can tell which change moved it.

Should mid-market products pick one tier and commit?

Not necessarily — but you need one page per tier with explicit numbers, rather than one page hedging across all of them. The brands that recovered fastest kept serving both ends and simply stopped describing themselves in a way that matched neither.

Does schema markup alone fix this?

No. Audience markup helped where body copy was ambiguous and did nothing where body copy contradicted it. It’s the last 10% of the job, not the first.

Can a competitor’s content cause my wrong size fit?

Yes, and it’s one of the harder versions to spot. In our nine-brand follow-up, both non-movers had a competitor comparison page ranking well that described them as built for large deployments — and the engines deferred to that over the brands’ own segment pages. Audit what ranks for "[your brand] alternatives" and "[your brand] vs" before concluding your on-site work failed.

Does the wrong size tier affect non-buying questions too?

It does. The same audience label leaks into adjacent question types — candidates asking whether you’re a good place to work get answers framed by whether the model thinks you’re a 40-person startup or a 4,000-person enterprise. Fixing the tier signal moves more than the shortlist.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →