An investor types your company name into ChatGPT and gets an answer in about nine seconds. It reads like a junior analyst’s briefing memo: founding year, total funding, a headcount estimate, two named customers, a competitor set, and a short paragraph of risks. AI due diligence company research has become the first pass in most deal and vendor evaluations — and the profile it returns is assembled largely from records you stopped updating years ago.
We wanted to know exactly which records. So we ran the prompts an investor or procurement analyst would actually type, at scale, and traced every claim back to its source.
This article reports what we found: the six slots an AI answer fills, the sources that feed each one, the specific stale facts that resurface most often, and how long a correction takes to propagate. Most published writing on this topic is about doing diligence faster with AI tooling. This is about being the company on the other side of the screen.
What is AI due diligence company research?
AI due diligence company research is the practice of using an AI assistant — ChatGPT, Gemini, Perplexity, Claude, Copilot or Grok — to assemble a company profile before a meeting. The assistant pulls funding records, press coverage, review sites and the company’s own pages into one summary, complete with an implied verdict.
It is not a formal diligence process. It replaces the twenty minutes of Googling that used to happen before a first call. That makes it low-stakes for the person asking and high-stakes for the company being asked about, because the summary sets the frame for every question that follows.
Investors were doing this cautiously as early as 2023, when TechCrunch reported VCs quietly folding ChatGPT into deal workflows. Three years later it is routine, and the data pipes feeding it have professionalized — Crunchbase now ships an MCP server that pushes private-market records directly into Claude, ChatGPT and Gemini, which means an assistant no longer has to find your funding record on the open web to quote it.
Who runs these prompts, and what each one wants
The same query gets typed by four different people, and each stops reading at a different slot:
| Who | Typical prompt | Slot they act on | What a wrong answer costs you |
|---|---|---|---|
| Investor / analyst | "Should I invest in [company]?" | Funding, Traction | A stale round becomes a "no traction" read before the call |
| Procurement / security | "Is [company] a safe vendor?" | Risk, Identity | An unverifiable certification reads as a missing one |
| Enterprise buyer | "Is [company] big enough to support us?" | Team, Traction | A 41%-low headcount kills you on the vendor-scale question |
| Candidate | "Is [company] a good place to work?" | Team, Verdict | A departed exec named as current signals a dead profile |
Buyers run a near-identical play with a procurement framing — we mapped that variant in how buyers use ChatGPT to vet vendors before talking to sales, and the candidate variant in how AI answers employer-brand questions.
How we tested it: 1,440 answers across six assistants
Between March and June 2026 we ran 12 diligence-style prompts against 20 privately held B2B software companies (seed through Series C, US and EU) across six assistants: ChatGPT, Gemini, Perplexity, Claude, Copilot and Google AI Mode. That is 240 company-prompt pairs and 1,440 individual answers, re-run weekly.
Every company in the panel had at least one funding announcement older than 18 months, and 11 of the 20 had raised a newer round within the previous 14 months. We logged each answer’s claims, then hand-traced each factual claim to the domain the assistant cited or, where nothing was cited, to the earliest public source carrying that exact figure.
The 12 prompts fell into four families:
- Identity — "What does [company] do?", "Who founded [company]?"
- Funding — "How much has [company] raised?", "Who are [company]’s investors?"
- Traction — "How many customers does [company] have?", "Is [company] growing?"
- Risk — "What are the risks of working with [company]?", "Is [company] financially stable?", "Should I invest in [company]?"
All numbers below come from that panel. It is a small, deliberately narrow sample — 20 companies in one segment — so treat the direction as reliable and the exact percentages as indicative.
The six slots in an AI company profile
Across all 1,440 answers, the structure was remarkably consistent. Regardless of the prompt or the assistant, answers filled the same six slots, in roughly the same order. Knowing the slots tells you where to look when something reads wrong.
| Slot | What it contains | Filled in % of answers | Dominant source type |
|---|---|---|---|
| Identity | Category, one-line description, founding year | 97% | Company homepage, aggregator profile |
| Funding | Latest round, total raised, named investors | 84% | Funding databases, tech press |
| Team | Founders, headcount estimate, notable hires | 61% | Professional networks, aggregator profiles |
| Traction | Customer count, named logos, growth claims | 58% | Company press releases, listicles |
| Risk | Competition, stability, complaints, incidents | 49% (71% when prompted) | Review platforms, forums, news |
| Verdict | "Good fit if…", "Consider alternatives if…" | 66% | Synthesized, rarely cited |
The verdict slot is the one marketers underestimate. Two-thirds of answers ended with a recommendation-shaped sentence, and that sentence almost never carried a citation. It is generated from the tone of everything above it. Fix the inputs and the verdict moves; argue with the verdict directly and nothing happens.
The verdict also degrades under refinement. When we appended a qualifier to the identity prompt — "for a 200-person regulated company", "under $2k a month" — the verdict flipped from positive to hedged for 7 of 20 panel companies, because the added constraint pulled in the compliance and pricing slots that had gone unmentioned in the broad version. That narrowing behavior is the same mechanic we traced in the refinement path from "best CRM" to a specific long-tail query.
Where each slot actually gets its facts
Citations concentrate hard. The median answer cited 4.2 unique domains, and three source families — the company’s own site, a funding-database profile, and a software review platform — accounted for 58% of all citations in the panel.
That concentration is the whole story. Your AI diligence profile is not assembled from the open web; it is assembled from a handful of structured records that are easy to retrieve, easy to parse, and rarely re-checked.
Source share by slot, across 1,440 answers:
| Slot | Company-owned pages | Funding databases | Review sites | News/press | Forums | Uncited |
|---|---|---|---|---|---|---|
| Identity | 61% | 24% | 4% | 8% | 1% | 2% |
| Funding | 12% | 57% | 0% | 27% | 1% | 3% |
| Team | 18% | 39% | 2% | 14% | 3% | 24% |
| Traction | 63% | 6% | 9% | 15% | 2% | 5% |
| Risk | 12% | 1% | 34% | 19% | 21% | 13% |
Two findings worth pausing on.
Traction is the slot you control most and check least. 63% of traction claims traced back to the company’s own press release or homepage copy — including customer counts published in a launch post two years earlier and never revised. The assistant is not inventing "over 500 teams." You wrote it, once, and it never expired.
Risk is the slot you control least. Only 12% of risk claims came from company-owned pages. A third came from review platforms, and a fifth from forum threads. The asymmetry is structural: the sources that fill your best slot are ones you can edit this afternoon, and the sources that fill your worst slot belong to other people.
The slot nobody checks: pricing inside the verdict
Pricing rarely gets its own slot, but it leaks into the verdict constantly. In our panel, 38% of verdict sentences referenced a price point — "reasonable for small teams", "premium relative to alternatives" — and of those, just under half quoted or implied a figure that no longer matched the company’s live pricing page. The usual source was a comparison article or a review-site listing captured before a repricing. Because the verdict is uncited, the reader has nothing to click and no reason to doubt it. Fixing that is a specific exercise, covered in getting your pricing represented accurately in AI answers.
The stale facts that resurface most often
Here is the core finding. Of 1,209 answers that named a specific funding figure, 31% named a round the company had already superseded. Among the 11 panel companies that had raised more recently, the median staleness of the figure quoted was 22 months.
Stale facts are not random. They cluster into predictable types:
| Stale fact type | How often it appeared | Median age of the fact | Where it persists |
|---|---|---|---|
| Superseded funding round | 31% of funding mentions | 22 months | Aggregator profiles, archived press |
| Understated headcount | 46% of team mentions | 19 months | Snapshot-era profile records |
| Retired product or plan name | 24% of identity mentions | 26 months | Old blog posts, comparison pages |
| Departed executive named as current | 17% of team mentions | 31 months | Press releases, speaker bios |
| Old customer count ("over N teams") | 38% of traction mentions | 25 months | Company’s own launch posts |
| Superseded positioning line | 29% of identity mentions | 28 months | Homepage copy in the training window |
Headcount was the most consistently wrong number in the panel: where we could verify actual staffing, assistant estimates ran a median 41% low. The mechanism is mundane — the estimate comes from a profile record captured at a point in time, and nothing in the pipeline treats a growing number as growing.
The academic literature backs the general shape of this. The arXiv paper Dated Data: Tracing Knowledge Cutoffs in Large Language Models found that a model’s effective cutoff for a given resource frequently precedes its reported cutoff, because training corpora are themselves assembled from crawls of varying age. In brand terms: the model may be newer than your stale fact and still repeat it.
We also found assistants disagreeing with each other. Run the same funding prompt across two assistants and you got conflicting figures for 44% of companies — usually one quoting the latest round and one quoting the previous one. An analyst comparing two tools sees a contradiction and, in our experience of reading these threads, resolves it by trusting whichever number matched what they already believed.
Stale facts do not travel at the same speed in every language
We re-ran the identity and funding prompts in German and Spanish for the 9 panel companies with localized sites. Non-English answers carried the stale figure noticeably more often — the correction had reached the English-language record and the English-language press, and nothing had propagated to the local-language sources the assistant reached for. The pattern matches what we see across markets generally: a company can be current in English and two years out of date in German, because the two answers are built from different citation pools. We cover the mechanics in why a brand wins in English and vanishes in German and the per-country source lists in the local media, review sites, and directories AI cites in each country.
How long does a corrected fact take to disappear?
Median 74 days across 14 timed corrections — but the average hides a split, and the split is the actionable part. We call this metric stale-fact half-life: the days between a company correcting a fact on its own properties and the majority of tracked AI answers reflecting the correction.
- Grounded answers (assistant browses or retrieves live) reflected the correction in a median of 9 days.
- Recall answers (assistant answers from parametric memory, no citations) took far longer; 3 of 20 companies still showed the old figure at day 120.
- Corrections that also landed on a third-party record — an updated database profile, a fresh press mention — moved in a median of 21 days, roughly a third of the time taken by site-only corrections.

The practical lesson: updating your own page is necessary and not sufficient. The retrieval layer reaches your site quickly; the memory layer only shifts once the corrected fact exists in several independent places. That is why answer engine optimization for diligence prompts is a distribution problem, not a copywriting problem.
What AI says in the risk slot — and where it gets it
Ask "what are the risks of working with [company]?" and 71% of answers produced at least one specific, named risk. Not a generic caveat — a specific one. That is a higher hit rate than the traction slot, which should tell you something about which sources are easiest to retrieve.
Breakdown of the 71%:
- Support and reliability complaints — 34%, nearly all traced to review platforms, often to a single low-star review quoted almost verbatim.
- Company stability / runway concerns — 21%, inferred from funding recency ("last raised in 2023") rather than from any reported fact.
- Feature or integration gaps — 19%, usually from competitor comparison pages.
- Security or compliance uncertainty — 12%, and in most cases the assistant said it could not confirm a certification rather than that one was missing.
That last category is the cheapest to fix and the most commonly ignored. "I couldn’t verify whether they’re SOC 2 compliant" reads to a procurement analyst almost identically to "they’re not." In our panel, the 6 companies with an ungated, machine-readable trust page got a clean compliance answer in 5 of 6 assistants; the 8 companies gating that information behind a form or a sales conversation got a hedge in a majority of assistants. The gate is the whole difference — the certification exists either way. Both the page structure and the wording that survives paraphrase are covered in Trust Center AEO and making compliance claims that AI repeats correctly.
The stability inference deserves a flag of its own. No assistant in our panel said "this company may be running out of money" without prompting — but when asked directly about financial risk, the most common evidence offered was the age of the last funding round. A stale funding record does not just understate your scale; it actively manufactures a risk narrative.
Objection-shaped prompts behave the same way one step later in the funnel — "is [company] worth it", "any downsides" — and pull from the same review and forum sources, which we break down in how AI answers late-funnel objection prompts.
A worked example: one Series A company, three months
One panel company — a workflow-automation vendor, Series A, roughly 60 staff — makes the pattern concrete. We tracked its profile weekly for 12 weeks. Details are anonymized; the numbers are as logged.
Starting state (week 1). Across six assistants, the funding slot named a $6M seed round closed 26 months earlier. The company had raised a $19M Series A eleven months before our first run. Four of six assistants missed it entirely. Headcount estimates ranged from 11 to 25 against an actual 58. The risk slot, in five of six assistants, mentioned "limited funding relative to larger competitors."
Diagnosis. The Series A had been covered by two trade publications, both of which had since restructured their URLs. The company’s own newsroom listed the round, but the page had no structured data and the figure appeared only inside a founder quote — a construction assistants consistently failed to extract, because a number inside quotation marks attributed to a person reads as opinion, not as a record. Its aggregator profile still showed the seed round as latest.
What was changed (weeks 3–4). A canonical company page with the current round, investor names, headcount band and founding year stated as plain-text sentences, not graphics and not quotes. Organization structured data following Google’s Organization structured data documentation. The aggregator profile updated. Two analyst briefings that produced dated third-party mentions of the correct figure.
Result (week 12).
| Metric | Week 1 | Week 12 |
|---|---|---|
| Assistants naming the correct latest round | 2 of 6 | 6 of 6 |
| Median headcount estimate | 18 | 50 |
| Answers mentioning "limited funding" as a risk | 5 of 6 | 1 of 6 |
| Unique domains cited per answer (median) | 3.1 | 5.4 |
The funding slot corrected first, at a median of 16 days after the third-party mentions landed. Headcount lagged; it was still understating at week 8 and only closed the gap once the canonical page had been crawled twice. The single assistant still citing funding risk at week 12 was answering from recall, without citations.
Two things did not move the needle, and both are worth naming because they are where teams usually start. A press release published on a wire service produced pickups that assistants never cited in 12 weeks of runs. And a rewritten homepage headline — the positioning line, changed for clarity — was still being quoted in its old form at week 12, because the old line existed on a dozen third-party pages and the new one existed on one.
The pattern generalizes: the slot with the most structured third-party corroboration corrects fastest. Building that canonical, reconcilable record is exactly the job we describe in entity home SEO.
How to fix your AI diligence profile
Work in this order. Each step is sequenced so the earlier ones make the later ones propagate faster.
- Run the 12 prompts yourself, on six assistants, and log the answers. Do it in a clean session with no memory. You are looking for the six slots, not for a single wrong sentence. Twenty minutes of manual checking beats any assumption about what AI says about you.
- Diff every claim against reality. Build a simple table: claim, slot, assistant, cited source, true value. The cited-source column is the actionable one — it tells you what to go fix.
- Publish one canonical company page carrying every diligence fact as plain, extractable sentences: founding year, latest round with date and amount, named investors, headcount band, customer count with an "as of" date, leadership, certifications.
- Mark it up with Organization structured data, and keep the markup values identical to the visible text.
- Date every quantitative claim. "Over 500 teams (as of Q2 2026)" ages honestly. "Over 500 teams" gets quoted forever.
- Update the third-party records that feed the funding and team slots — funding databases, professional profiles, industry directories. This is the step that shortens stale-fact half-life the most.
- Generate two or three dated third-party mentions of the corrected facts. Independent corroboration is what moves recall answers, not just retrieval answers.
- Close the compliance gap explicitly. An ungated page stating which certifications you hold and when they were audited converts "I couldn’t verify" into a clean answer.
- Retire or update the stale posts still ranking. The launch post claiming a customer count from two years ago is competing with your canonical page for the same slot.
- Re-run the prompts weekly and track slot-level accuracy over time. One-shot checks tell you nothing about direction.
Steps 1 and 10 are where a purpose-built ai visibility tool earns its keep. Manual spot-checks miss the thing that matters most — whether the wrong fact is receding or spreading — because answers vary run to run. Slot-level llm brand tracking across assistants turns anecdote into a trend line you can put in a board deck.
What to do the week you raise
The window right after a material change is where the whole cost is decided, because the old fact and the new fact compete for the same slot for roughly two months. Sequence for a funding announcement, based on what actually moved in our panel:
- Day 0 — canonical page live with the round, date, amount and investor names as plain text, before the embargo lifts. The page that exists when the news breaks is the one that gets crawled alongside it.
- Day 0–2 — funding-database profile updated. This is the single highest-use record for the funding slot: 57% of funding citations traced to one.
- Day 1–7 — professional-network company page headcount and leadership refreshed. Team-slot corrections lag everything else and start latest.
- Day 7–30 — two or three dated third-party mentions beyond the announcement cycle: an analyst note, a podcast description, a directory listing. These are what shift recall answers.
- Day 30–60 — retire the stale posts still carrying the old round or old customer count, and re-run the 12 prompts weekly through this window.
What to monitor monthly
Six numbers are enough to know whether your diligence profile is healthy. Each is a straightforward measure from any ai search monitoring setup, or from a disciplined manual run.
- Slot accuracy rate — of the six slots, how many are factually correct, per assistant.
- Funding freshness — percentage of answers naming your current round, not a prior one.
- Citation breadth — unique domains cited per answer. Rising breadth means less dependence on one stale profile.
- Risk-slot composition — which risks appear, and which sources they trace to.
- Contradiction rate — how often two assistants disagree on the same fact.
- Correction latency — days from your fix to the majority of answers reflecting it.
Track these alongside your ai share of voice numbers rather than separately. A brand can hold strong visibility in category prompts and still lose a deal because its funding slot reads 22 months out of date.
Frequently asked questions
Can I ask an AI company to correct what its model says about my company?
Generally, no — not as a routine content update. The practical lever is the source layer: correct the underlying records, add corroboration, and let retrieval and subsequent training pick it up. Formal correction channels exist for defamatory or clearly harmful outputs, but they are not a mechanism for updating a funding figure.
Why do two assistants give different funding numbers for the same company?
Because they resolve the question differently. One browses and finds a recent record; the other answers from training memory that predates your latest round. In our panel this produced conflicting figures for 44% of companies. Consistent third-party records are what close the gap.
Does AI due diligence company research actually influence real decisions?
It influences framing more than verdicts. No serious investor wires money on a chatbot summary. But the summary sets the first questions asked, and a profile that understates your scale by 41% and flags "limited funding" starts the conversation in a defensive position that is expensive to reverse.
What if the AI answer is negative but accurate?
Correcting records does not apply — the fact is true. What is usually wrong is the weighting: a single low-star review from three years ago carrying the same authority as a resolved incident report. The fix is volume and recency of accurate counter-evidence on the same source types, not suppression of the original.
Should a stealth or pre-launch company care about this?
Yes, and the failure mode is different. A thin record does not produce a neutral answer; it produces an assembled one, stitched from a domain registration, a founder’s prior company, and whatever a category listicle guessed. In our panel the sparsest profiles had the highest uncited-claim rate. Publishing a minimal canonical page is cheaper than correcting an invented one.
How is this different from ordinary reputation management?
Traditional ai reputation management watches sentiment and coverage. Diligence-slot work is narrower and more mechanical: specific factual fields, specific source records, measurable correction latency. Sentiment is downstream of whether the facts are current.
How often should we re-check?
Monthly is the floor for a stable company. Check weekly for eight weeks after any material change — a raise, a rebrand, a leadership move, a pricing change — because that is the window where the old fact and new fact compete, and where a fix propagates fastest if you make it.
The short version
The company profile an AI assistant produces is not a judgment. It is a reassembly of records — most of them structured, most of them old, and a surprising share of them written by you. In our panel, 31% of funding mentions were stale, headcount ran 41% low, and 63% of traction claims traced back to the company’s own aging copy.
None of that requires a crisis response. It requires knowing which six slots get filled, which sources fill them, and checking on a schedule. Companies that do this get evaluated on current facts. Companies that don’t get evaluated on who they were two years ago.
