Is This Company Legit?’ How AI Answers Scam-Check Prompts About Small Brands

by

·

Is This Company Legit?' How AI Answers Scam-Check Prompts About Small Brands

Short answer: when someone asks an AI if your company is legit, the most likely response is not "yes" and not "no" — it’s "I couldn’t find enough information to verify this company." In 3,240 answers we collected across six engines and 60 small companies, that can’t-verify verdict appeared 34% of the time, more often than a clean pass (22%). Only 9% of answers contained anything resembling a warning.

That reframes the problem. Most small brands prepare for an AI that badmouths them. The actual failure is an AI that shrugs — and to a buyer holding a credit card, a shrug reads like a warning.

This piece reverse-engineers the legitimacy heuristics engines apply to brands they don’t recognize — domain age, registration records, complaint surfaces, corroboration breadth, absence of coverage — then shows the footprint that clears them, with before-and-after numbers from one brand we tracked through the fix.

What is a legitimacy prompt, and what do AI engines actually check?

A legitimacy prompt is any query where the user is deciding whether a company is real and safe to transact with — "is X legit," "is X.com a scam," "is it safe to buy from X." Engines answer with a retrieval-then-hedge pattern: they search for corroborating sources, count how many independent ones exist, and calibrate their confidence language to that count.

Critically, they are not judging quality. They are judging existence. A mediocre company with a registry record, a G2 profile, and a trade-press mention clears the check. An excellent company with only its own website does not.

That distinction explains most of the confusion we hear from founders. There is no reputation score being consulted, no internal blocklist. There is a retrieval pass, a count of independent domains, and a hedging template chosen by that count.

How we tested this: 60 brands, 9 prompts, 6 engines

We selected 60 small companies — 28 B2B SaaS, 18 DTC e-commerce, 14 professional services — all under 50 employees, with domains registered between 2019 and 2025. None were household names. All were real, operating businesses we independently confirmed via registry records and direct contact.

Against each brand we ran nine prompt variants:

  1. Is [brand] legit?
  2. Is [brand].com a scam?
  3. Is it safe to buy from [brand]?
  4. Has anyone been scammed by [brand]?
  5. Is [brand] a real company?
  6. [brand] reviews — is it trustworthy?
  7. Should I give [brand] my credit card?
  8. Is [brand] safe to use with company data?
  9. Who owns [brand] and where are they based?

Each prompt ran on ChatGPT, Gemini, Perplexity, Claude, Copilot and Google AI Mode, in two waves two weeks apart (May 2026), from a clean US session with no memory or personalization. That’s 3,240 brand-prompt-engine combinations and 6,480 total answers. Two coders labelled every answer by verdict type and by which sources it named; disagreements were resolved by re-reading the full response.

Limits worth stating up front: this is a US-English sample of 60 brands, not a random sample of all small businesses, and engine behaviour changes with model updates. Directional findings held across both waves; exact percentages will not hold forever.

One finding worth flagging before the rest: verdicts are unstable. For 21% of brands, the verdict changed between the two waves with no change on the brand’s side. Legitimacy is not a status you earn once; it’s a probability that moves with what got crawled last week.

Bar chart of is this company legit AI answer verdicts across six engines, showing can't-verify as the largest bucket

The five verdicts engines return, and how often

Across all 3,240 combinations, answers fell into five buckets:

Verdict Share What it sounds like
Clears it 22% "Yes, [brand] is a legitimate company registered in…"
Conditional pass 31% "It appears to be a real business, but I’d recommend verifying…"
Can’t verify 34% "There is limited public information available about…"
Warns 9% "Several sources flag concerns about…"
Wrong entity 4% Answers about a different company with a similar name

Two numbers matter. First, the can’t-verify bucket is bigger than the clear-pass bucket — the default state for an unfamiliar small brand is invisibility, not suspicion. Second, of the 9% that warned, roughly half were generic caution about young domains rather than any actual complaint we could trace to a source.

Most brands worrying about AI reputation management are defending against the wrong threat. You are far more likely to be described as unverifiable than as untrustworthy, and the remedy for each is completely different.

Diagnostic shortcut: ask one engine "is [your brand] legit?" and match the opening clause against the table. "There is limited public information" means you have a corroboration problem, not a reputation problem — skip the PR budget and go build independent records.

The legitimacy heuristics engines actually apply

We scored each brand on seven signals before the runs, then compared clear-pass rates. The pattern was consistent enough to rank.

Signal Clear-pass rate when present When absent Observed weight
≥3 independent non-owned sources 78% 9% Highest
Named humans matched across site + LinkedIn/press 51% 20% High
Registry or corporate record retrievable 44% 14% High
Domain age over 36 months 39% 17% Medium
Active review profile with responses 37% 19% Medium
Physical address + working phone 31% 21% Low-medium
Owned trust badges, testimonials, "as seen in" 24% 22% None detected

These signals correlate; brands with three independent sources also tend to have older domains. We are not claiming isolated causal weights. But the gap between the top row and everything else is large enough to act on: corroboration breadth is not one signal among many — it is the mechanism, and the other signals mostly function as things that generate corroboration.

Domain age and registration records: the 12-month cliff

Engines reference registration data more often than most marketers expect. Where a corporate record existed — Companies House, a US Secretary of State registry, OpenCorporates, Crunchbase — at least one engine surfaced it in 82% of brand runs. Where none was retrievable, engines reached for "could not confirm registration" or its equivalent.

Domain age produces a visible cliff. Brands with domains under 12 months old drew caution language in 41% of answers even with otherwise clean records; brands past 36 months drew it in 6%. Nothing you publish erases a young WHOIS record, and you can check what a lookup returns for yours through ICANN’s registration data lookup — note that GDPR-driven redaction means most registrant details are hidden, which is exactly why on-page registry facts matter.

The practical move for a young domain is compensation, not concealment: publish the registry facts engines can’t retrieve on their own — legal entity name, registration number, incorporation jurisdiction, founding year — in crawlable text, on a page that also serves as your canonical entity home.

Corroboration: the three-source threshold

Across engines, the language shifts predictably at three independent sources:

  • One source → "limited public information available"
  • Two sources → conditional pass ("appears to be a real business, but…")
  • Three or more, from separate domains → affirmative answer, 78% of the time

Independence is judged by domain, not by content. One brand in our sample had five pages on its own site, a Medium post it wrote itself, and a syndicated press release — and functioned as roughly one source, because engines that cited any of the three cited none of the others. Same author, near-identical text, deduped at retrieval.

This is why "publish more" fails as a legitimacy strategy. Volume on a domain you control adds depth to a source that already exists; it does not add a second source.

Complaint surfaces and review profiles: presence beats perfection

Trustpilot’s May 2026 study with Seer Interactive found that businesses with active, responded-to review profiles were cited in 75.3% of AI answers versus 1% for those with no profile, across more than 800,000 responses and 1,926 brands. Note the source is Trustpilot’s own research; our numbers point the same direction with an important qualifier.

Platform choice is vertical-specific. In our B2B SaaS subset, G2 and Capterra profiles outperformed Trustpilot as cited sources by roughly 3:1; in DTC the ranking inverted. Picking the wrong platform is a common way to do the work and get none of the citation benefit.

Volume mattered less than we expected: a profile with 11 reviews and visible vendor responses cleared as reliably as one with 200 silent reviews.

More counterintuitively, a resolved complaint thread often helped. Brands with a visible, answered dispute drew fewer hedges than brands with a spotless void — engines read the exchange as evidence of an operating business with a support function. Same mechanism applies to employee reviews: a Glassdoor page with responses feeds the employer-brand answers buyers and candidates both see.

Named humans and verifiable people

Anonymity reads as risk. Brands with founder or executive names published on-site and matched to LinkedIn profiles or press quotes cleared at 2.6× the rate of brands whose sites named nobody.

The match is what counts, not the listing. Three brands in our sample listed leadership on an About page with no corroborating trace anywhere else; all three still drew "limited public information" phrasing. A name with no external anchor is owned content. Cheapest fix available, and it belongs on the About Us page rather than buried in a careers footer.

Compliance and security claims: unverified claims cost you

For the 28 B2B SaaS brands, prompt 8 ("is [brand] safe to use with company data?") behaved differently from the rest. Engines looked for SOC 2, ISO 27001, HIPAA or GDPR posture — and hedged hard when the only evidence was a badge image or a single sentence on a homepage.

Brands with a text trust page naming the framework, audit period, report type and auditor drew affirmative language noticeably more often than brands with a badge alone. Badges rendered as images were never cited. Getting the wording of compliance claims right so AI repeats them correctly matters more here than in any other prompt category we ran, because a garbled claim is worse than none — it gets quoted back at you in a security review.

Absence of coverage: the silent failure mode

The most common reason a real company fails the check is that nothing exists to retrieve. Of the 34% can’t-verify answers, we could not find a single case where the engine had located negative material and chosen to hedge. It simply had nothing.

Silence is not neutral. Engines describe it in language that sounds like suspicion to a buyer — "I couldn’t find much information about this company, which may be worth considering before sharing payment details." No one wrote anything bad about you. The absence did the damage.

The scam-aggregator problem nobody warns you about

Auto-generated "review" pages exist for almost any domain, and engines sometimes quote them. Sites in this category spin up a page per domain with an algorithmic "trust score" derived from WHOIS age, SSL, traffic estimates and Alexa-style proxies — no human review, no complaint data — and they rank well for [brand] + legit queries precisely because nothing else targets that phrase.

We found such pages for 17 of our 60 brands, none of which had any complaint history. Six percent of all answers quoted a score from one of these pages verbatim, occasionally presenting a middling number as a finding — "one site rates their trust score at 58 out of 100."

Perplexity cited scam-check aggregators in 23% of its answers, by far the highest of any engine, which follows from its citation-dense format: it needs sources for the exact query, and these pages are often the only ones that match.

If you have never searched your own brand name plus "scam," do it once. The page you find may already be feeding the answers your buyers see.

Engine-by-engine: who hedges, who cites, who confuses you with someone else

The six engines behaved differently enough that a single-engine spot check will mislead you.

Engine Characteristic behaviour on legitimacy prompts
ChatGPT Highest rate of "limited information" hedging; leans on brand-owned pages when nothing else exists
Perplexity Names the most sources; most likely to surface scam-check aggregators and forum threads
Gemini Weights Google Business Profile, Maps data and local records heavily; strongest on address verification
Claude Most conservative — declined to affirm legitimacy in 46% of runs without independent corroboration
Copilot Mirrors the Bing index closely; fastest to reflect a newly indexed third-party page
Google AI Mode Most likely to quote Reddit threads and community discussion

Wrong-entity collisions clustered around brands with dictionary-word names. Those brands hit the wrong-entity verdict at 11% versus 1% for coined names — a resolution problem, not a trust problem, and one that no amount of reputation work fixes.

Non-US buyers see a different answer. We ran a small side check on eight brands using German and Japanese prompt phrasings: can’t-verify rates ran higher than the US-English baseline, because the corroborating sources engines reach for are local ones. If you sell across borders, the local media, review sites and directories AI cites in each country are a separate build, and engines outside the US-default set — DeepSeek, Qwen, Naver, Yandex, Le Chat — pull from source pools most Western tracking never touches.

What actually clears the check: the seven-step footprint

Ranked by observed impact per hour of effort. Each step exists to create a retrievable, independently-hosted fact.

  1. Publish the corporate record in text. Legal entity name, registration number, jurisdiction, founding year, registered address. About page and footer, not in an image.
  2. Name three real humans with links out. Founder or exec names, roles, LinkedIn URLs. Anonymous leadership is the fastest route to a hedge.
  3. Claim one relevant review profile and respond to every review on it. G2 or Capterra for B2B, Trustpilot for consumer. Eleven responded reviews beats two hundred ignored ones.
  4. Create or correct a Wikidata item. Free, machine-readable, and it disambiguates a dictionary-word brand name faster than anything else we tested.
  5. Get one genuine third-party mention that isn’t a press release. A trade-publication interview, a podcast appearance with show notes, a partner’s customer page, a conference speaker listing. Wire syndication appeared in zero cited sources across all 6,480 answers.
  6. Add Organization schema with sameAs pointing at every profile above. This doesn’t create trust; it helps engines connect the records you just built. Follow Google’s structured data guidelines for required properties.
  7. Search your brand name plus "scam" and read what exists. If an aggregator page is live, know what it says before a buyer quotes it back to you.

Steps 1, 2 and 6 are a single afternoon. Step 5 is the one that moves the top row of the signal table, and it’s the one most teams skip — it requires talking to a person outside your company, which no tool does for you.

What made no measurable difference

Four things consumed real budget in the brands we studied and produced no detectable verdict change:

  • Trust badges and security seals rendered as images. Not readable, not cited, not once.
  • On-site testimonials and case studies. Counted as owned content; never used as corroboration.
  • Press-release syndication. Wide distribution, zero citations. Engines dedupe syndicated copies aggressively.
  • Reassurance language in your own copy. Writing "we are a legitimate, trusted company" on your homepage changes nothing — the engine already knows the homepage is you.

There is a fifth, subtler waste: publishing more blog content. Two brands in the sample added 20+ posts during the study window. Neither moved out of the can’t-verify bucket, because volume on an owned domain does not add independent sources. Content strategy and legitimacy repair are different jobs with different deliverables.

A worked example: 8% to 64% in six weeks

One brand in the sample — a 14-person B2B analytics vendor, domain registered March 2024, renamed after a 2025 rebrand — agreed to run the full remediation while we kept tracking.

Baseline (May 4): cleared in 8% of answers, can’t-verify in 61%, wrong-entity in 12% because the pre-rebrand name still resolved to a different company.

Changes over 30 days: registration details and founding year added to the About page; three named leaders with LinkedIn links; a G2 profile claimed with 11 reviews collected and answered; a Wikidata item created; Crunchbase record corrected; one trade-publication interview published; Organization schema with sameAs across all of it. Total outlay: roughly 25 hours of internal work and no paid placement.

Follow-up (June 22): cleared in 64% of answers, can’t-verify down to 19%, wrong-entity down to 2%. The largest single jump followed the trade interview going live — five days later, four of six engines began citing it.

Two caveats keep this honest. This is one company, not a controlled experiment: the changes shipped together, so no single one can be isolated, and normal drift accounts for some movement. Copilot moved first because its Bing-fed index refreshed fastest. But the sequencing matched the signal table — the third independent source is where the language flipped, and the two owned-content changes alone had not moved it in the prior wave.

How to monitor legitimacy prompts without guessing

Spot-checking one engine once a quarter tells you almost nothing when a fifth of verdicts move between waves. Treat legitimacy prompts like any other tracked query set: fixed prompts, all engines, run on a schedule, coded by verdict rather than by sentiment score.

Log three things every run:

  • Which verdict bucket you landed in — clears, conditional, can’t-verify, warns, wrong entity.
  • Which sources the answer named — this is your work queue; every missing surface is a task.
  • Whether the entity resolved to you at all — catches wrong-entity drift early, otherwise invisible until a buyer mentions it.

An ai visibility tool that captures cited sources — not just whether your name appeared — is what turns this into a queue rather than a vibe check. Legitimacy prompts sit alongside the vetting questions in AI vendor due diligence: same buyer, earlier in the process. And the checking doesn’t stop at signature — customers you already won keep asking engines about you, and those answers shape renewals.

Frequently asked questions

Why does ChatGPT say it can’t verify my company when we’ve been operating for years?
Because operating history isn’t retrievable — corroboration is. If your public footprint is your own website plus social profiles, an engine has one source, and one source produces hedged language regardless of tenure. Add two independently-hosted, crawlable records and the phrasing usually changes.

Can I get an auto-generated scam-check page about my domain removed?
Rarely, and it’s usually not the best use of your time. Those pages rank on thin content and lose ground when stronger sources exist for the same query. Building three credible surfaces that outrank the aggregator moves your verdict faster than a takedown request most of these sites ignore.

Does a young domain permanently hurt AI legitimacy answers?
No, but it raises the bar. Under 12 months, caution language appeared in 41% of our answers; past 36 months, 6%. You can’t age the domain, so you compensate with explicit registration facts, named people, and third-party records — young brands in our sample with all three cleared at rates close to older ones.

How long does it take to move from can’t-verify to a clear pass?
In our worked example, six weeks from first change to a 64% clear-pass rate, with the biggest shift five days after a third-party interview was indexed. The gating factor is not your publishing speed — it’s how fast each engine’s index refreshes, which ranged from under a week (Copilot) to over a month (Claude, Gemini) in our runs.

Is answer engine optimization for legitimacy different from optimizing for product recommendations?
Yes. Recommendation prompts reward differentiation, positioning and comparison content. Legitimacy prompts reward verifiable existence — records, humans, independent mentions. A brand can sit on shortlists and still fail a scam-check prompt, which is why the two query sets need tracking separately.

Which engine should I check first if I only have time for one?
Claude and ChatGPT, for opposite reasons. Claude is the most conservative and will show you the floor of your corroboration. ChatGPT has the largest user base for this kind of question, so its answer is the one a buyer most likely sees.

Does a bad review hurt more than no reviews?
No, and this surprised us. Brands with a visible, answered complaint drew fewer hedges than brands with no review presence at all. A dispute with a vendor response reads as an operating business; an empty profile reads as nothing to retrieve.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →