Multilingual AI Search Visibility: Winning in English, Invisible in German

by

·

Side-by-side ChatGPT answers to the same question in English and German, showing multilingual AI search visibility gaps in the recommended brand list

Multilingual AI search visibility collapses at the language border, not the national border. Across 4,000 AI answers we logged in June 2026, twelve B2B SaaS brands averaged a 38.2% mention rate when the prompt was written in English — and 17.3% when a native speaker asked the same buying question in German.

Same buyer. Same product. Same model, same week. The only variable was the language the question was typed in.

That matters because the standard fix — VPNs, country pages, geo-targeting — addresses the variable that barely moved in our test. Below: what we measured, the four diagnosable reasons a brand disappears in a language, and a nine-week case where a 40-point gap closed to 19.

What is multilingual AI search visibility?

Multilingual AI search visibility is how often, how prominently, and how accurately an AI assistant names your brand when the same buying question is asked in different languages. It is measured per language, per engine and per prompt — never as one global number — because each language draws on a partly separate pool of retrievable sources.

The practical consequence: a brand can hold a top-three slot in English answers on ChatGPT and be entirely absent from the German answer to the same question, on the same model, on the same day. One brand tracker running one language is a tracker covering one market.

Side-by-side ChatGPT answers to the same question in English and German, showing multilingual AI search visibility gaps in the recommended brand list

Why one model recommends different brands in each language

Models do not translate their opinion of your brand. For a commercial "best X for Y" question, the assistant retrieves in the language of the prompt, then writes an answer grounded in what it pulled. The brand list is assembled from that retrieval — so the answer reflects the language-native corpus, not a translated version of the English one.

Three forces compound:

  • Volume asymmetry. English-language software commentary — review sites, Reddit threads, comparison posts, trade press — outweighs German-language equivalents by roughly an order of magnitude in most B2B categories. Thinner corpus, fewer candidate brands, more weight on each surviving source.
  • Different gatekeepers. German buyers read OMR Reviews and Trusted; French buyers read Appvizer; Japanese buyers read ITreview and BOXIL. If those specific properties don't carry your brand, the in-language corpus has nothing to retrieve.
  • Independent corroboration. A single vendor-owned page rarely produces a recommendation. Agreement across independent sources does — the same mechanism behind off-site signals in AI search, except it has to be re-earned in every language.

This is also why gains do not transfer. The corpora are largely independent, which cuts both directions.

How the five-language test was run

We ran a controlled prompt set rather than a source-citation crawl, because the commercial question is does the brand get recommended? — not does the domain get linked?

Test design (June 8–26, 2026):

Parameter Value
Brands tracked 12 B2B SaaS brands, 3 categories (HR software, customer support, product analytics)
Prompts 40 buyer-intent prompts ("best X for Y", "alternatives to Z", "which X handles GDPR")
Languages English, Spanish, French, German, Japanese
Engines ChatGPT, Google AI Overviews, Perplexity, Gemini
Repeat runs 5 per prompt / language / engine
Total answers scored 4,000

Every prompt was rewritten by a native speaker, not machine-translated. A machine-translated prompt tests your translation memory, not the market.

Mention rate = the share of the 20 answer slots per prompt-language (4 engines × 5 runs) in which the brand appeared anywhere in the response text. Five runs is the floor, not a luxury: single-run readings on these engines swing enough to invent trends that aren't there, which is why how many repeat runs an AI visibility number needs is a prerequisite question rather than a footnote.

What the results showed

The gap widened consistently with distance from the English corpus:

Prompt language Mean mention rate Change vs. English Brands losing >half their English rate
English 38.2%
Spanish 24.1% −37% 3 of 12
French 20.9% −45% 5 of 12
German 17.3% −55% 7 of 12
Japanese 9.4% −75% 10 of 12

Three findings stood out beyond the averages.

Absence was common, not marginal. Four of the twelve brands never appeared once in any Japanese answer across 200 scored slots. Two never appeared in German. These were not small companies — all twelve had eight-figure ARR and functioning localized websites.

The vacancy was filled locally. In German prompts, a domestic vendor took the first-named position in 61% of answers, against 12% for the same prompts in English. AI share of voice did not evaporate; it transferred to whichever brand the German-language corpus discusses.

Translated prompts flatter you. We also ran literal translations of the English prompt list alongside the native-phrased set. The literal-translation German prompts returned 23.4% mention rate versus 17.3% for native phrasing — 6.1 points of pure measurement error, in the optimistic direction, for any team that builds its multilingual prompt set with a translation tool.

The gap is not the same on every engine

Prompt language ChatGPT Google AI Overviews Perplexity Gemini
English 41.0% 34.5% 40.2% 37.1%
Spanish 22.4% 28.9% 23.6% 21.5%
French 19.2% 25.7% 20.1% 18.6%
German 15.1% 21.8% 16.4% 15.9%
Japanese 7.2% 13.5% 8.8% 8.1%

Read the retention, not the absolute numbers. Google AI Overviews retained 63% of its English mention rate in German; ChatGPT retained 37%. Google's surfaces start lower in English and degrade least — decades of geo-localized, language-aware retrieval carry over into AI features built on that index. The pure chat assistants start higher and fall furthest, because they lean almost entirely on the prompt language.

Practical read: if your only non-English tracking is on Google surfaces, you are measuring your best case.

Language or location: which one actually moves the answer?

Prompt language moved brand mentions roughly 16× more than request location in our test. We ran the same 40 German prompts twice — once from a US (Virginia) egress, once from a Frankfurt egress — holding engine, run count and wording constant.

Condition Mean mention rate Delta
English prompt, US location 38.2% baseline
German prompt, US location 17.3% −20.9 pts
German prompt, Frankfurt location 18.6% +1.3 pts vs. US

Median run-to-run noise at five runs was ±2.8 points. The 1.3-point location effect sits inside the noise band; the 20.9-point language effect sits far outside it.

The location effect was not uniform across engines, and reporting it as zero would be wrong:

  • Google AI Overviews: +4.4 pts from Frankfurt — the only engine clearing the noise band
  • Gemini: +2.1 pts
  • Perplexity: +0.6 pts
  • ChatGPT: +0.2 pts

Weglot's analysis of 1.3 million AI citations reported the reverse emphasis — that testing from a Mexico City connection raised the share of Spanish-language sources cited from 32% to 63%. Both results hold, and reading them together is more useful than picking one: location strongly changes which sources get cited; language strongly changes which brands get named. In our runs, a German answer served from Frankfurt cited more .de domains than the same answer served from Virginia — and still recommended the same brands.

Language and geography therefore need separate tracking axes. Monitoring recommendations across cities solves a different problem than the one described here; swapping one for the other produces a dashboard that looks complete and measures the wrong thing.

Four reasons a brand vanishes in another language

Every zero-mention case in our test resolved to one of four diagnosable causes. Naming the right one matters, because the fixes are not interchangeable.

1. Corpus absence

The in-language corpus contains no substantive mention of the brand. The model isn't ignoring you; it has nothing to retrieve.
Signal: ask the engine directly in-language — "Was ist [brand]?" A hedge or a hallucinated answer confirms it. This was the cause in 5 of our 12 brands' German gaps.

2. Entity fragmentation

The brand exists in the corpus under a name variant the model does not resolve to a single entity — a transliteration, a local legal name, a "GmbH" suffix, or a script change. The knowledge exists but is split across two half-entities, neither strong enough to surface. This failure mode dominated in Japanese, where katakana and Latin-script renderings of the same brand behaved as separate entities in 6 of 12 cases.

3. Translation-shell content

Localized pages exist, but they are machine-translated mirrors carrying no independent citations, no local customer names and no market-specific claims. They are retrievable and worthless as corroboration. Brands with natively written German pages held 2.3× the German mention rate of brands with translated mirrors at comparable page counts.

4. Category-term mismatch

Your content targets the English category label; local buyers use a different concept. German buyers ask about Personalverwaltung and Zeiterfassung, not "HR platform" — so retrieval never reaches your cluster, however well-optimized that cluster is. This is a prompt-set problem before it is a content problem, and it is invisible if your tracked prompts are translations of the English list.

Failure mode Fastest diagnostic Typical time to move
Corpus absence In-language "what is [brand]?" returns nothing usable 6–12 weeks
Entity fragmentation Compare answers for Latin-script vs. local-script name 4–8 weeks
Translation-shell content Check whether cited pages are MT mirrors 8–16 weeks
Category-term mismatch Compare your prompt list to native-speaker phrasing 1–3 weeks

How to diagnose your own language gap in one afternoon

  1. Pick one non-English market with real revenue attached. Not five. Diagnosis is cheap; remediation is not.
  2. Take your ten highest-value English prompts and have a native speaker rewrite them — rewrite, not translate. Ask what they would actually type.
  3. Run both sets five times each on at least two engines, from one location, on one day. Holding location constant is the whole point.
  4. Score three metrics per language: mention rate, first-named share, and descriptor accuracy (does the model describe your category correctly?). Descriptor accuracy usually breaks before mentions do.
  5. For every prompt where you're absent in-language, capture the cited sources. That list is your remediation backlog — it names the exact properties the model trusts in that market.
  6. Re-run weekly with English as your control. If English moves too, your intervention wasn't the cause; the logic behind running a holdout test when you only have one brand applies directly.

Ten prompts × 2 languages × 2 engines × 5 runs = 200 answers. That is an afternoon of automated collection and about an hour of reading.

Building a native prompt set — the step teams skip

The prompt list is where most multilingual programmes go wrong, because a translated list is quietly answerable by your English-derived content.

English prompt Literal translation (what most teams track) What buyers actually type
best help desk software for small teams beste Helpdesk-Software für kleine Teams Ticketsystem für kleine Unternehmen
HR platform for remote teams HR-Plattform für Remote-Teams Personalverwaltung Software Remote-Mitarbeiter
best CRM for startups スタートアップ向けの最適なCRM 中小企業 CRM 比較 おすすめ

Sit with a native-speaking salesperson or customer for 30 minutes and write down the nouns they use for your category. That list is worth more than any keyword tool export, and it is the input to everything downstream.

Dashboard comparing mention rate, first-named position and descriptor accuracy across five prompt languages

The 90-day remediation sequence

Order matters, because each step makes the next cheaper.

  • Weeks 1–2 — fix the measurement. Native prompt set, five runs, English held as control. Diagnose which of the four failure modes you have. Do not commission content yet: a category-term mismatch changes what content you would commission.
  • Weeks 2–4 — fix the entity. One canonical in-language name, used identically on your site, in review-site profiles, in press, and on Wikidata if you qualify. Cheap, fast, and it makes every later mention count toward the same entity instead of splitting.
  • Weeks 3–8 — earn independent in-language sources. Review-site profiles with real verified reviews, one or two trade-press placements, one named local customer. This is the slow, load-bearing part.
  • Weeks 6–12 — rewrite the highest-intent pages natively. Not the whole site. Pricing, your top two category pages, integrations. Written by someone who sells in that market.
  • Ongoing — re-run weekly. Watch cited-source coverage first; it moves before mention rate.

Case study: closing a 40-point German gap in nine weeks

One tracked brand — a product analytics platform, anonymized as Brand C — opened the window at a 46% English mention rate and a 6% German mention rate. Worse, German answers described it as an "E-Mail-Marketing-Tool" in 3 of 5 ChatGPT runs. Wrong category, confidently stated.

Diagnosis: translation-shell content plus corpus absence. The German site was a complete MT mirror. German-language third-party mentions: zero.

What the team changed over nine weeks:

  • Three German review-site profiles built out, each to 10+ verified reviews
  • Two German trade-press mentions earned (product launch coverage, one analyst roundup)
  • Four pages rewritten natively in German — pricing, the two highest-intent category pages, integrations
  • One German-language case study with a named, referenceable DACH customer

Result at week nine: German mention rate 6% → 27%. English mention rate 46% → 47% — flat, which is what makes the German movement attributable. Miscategorization hit 0 of 5 runs by week six, ahead of the mention gain: descriptor accuracy repairs faster than recommendation share.

Not everything helped. Correcting hreflang annotations produced no measurable change, because they were already correct — though Google's guidance on signalling localized versions remains table stakes for the classic index. Rewriting German meta descriptions produced no measurable change either. The movement came from off-site sources in German, not from on-site tuning.

What consistently fails to move the number

  • Bulk machine translation of the whole site. Volume of translated pages showed no correlation with mention rate in our set; number of independent in-language sources did.
  • Translating your English prompt list. It inflates your reading by ~6 points and points remediation at questions nobody asks.
  • Tracking only US-centric engines. Japanese and Korean gaps look materially different once locally dominant answer surfaces are in the panel.
  • Reading a single run as a result. At five runs our noise band was ±2.8 points; at one run, roughly half the "improvements" a team would report are sampling artifacts.
  • Buying a ccTLD and waiting. Domain geography changed nothing measurable in our data. The corpus that mentions you is the lever.

How to report this to a budget holder

Report one number per market, never one global average — an average hides exactly the failure you are trying to fund fixing.

Metric What it answers Report cadence
Mention rate by language Do we appear at all? Weekly
First-named share Do we appear first? Weekly
AI share of voice vs. top local rival Are we losing to a domestic incumbent? Monthly
Descriptor accuracy Is the model describing us correctly? Monthly
Cited-source coverage How many trusted local properties mention us? Monthly

Cited-source coverage is the leading indicator. In our data it moved four to six weeks before mention rate — the only metric on the list you can show an executive as evidence that a nine-week programme is working before the headline number confirms it.

For teams building the full per-market programme rather than a one-off diagnostic, the sequencing is laid out in the multilingual AEO playbook for winning abroad.

Frequently asked questions

Does using a VPN change what AI assistants recommend?

Marginally, and mostly on Google's surfaces. Switching from a US to a German egress moved German mention rate 1.3 points overall — inside the ±2.8-point noise band — while Google AI Overviews alone moved 4.4 points. Changing the prompt language moved it 20.9 points. Make language your primary variable and location your secondary one.

Do I need separate content per language, or is translation enough?

Translation gets you retrievable; it rarely gets you recommended. Brands with natively written pages held 2.3× the German mention rate of brands with MT mirrors at comparable page counts. The bigger lever is off-site: independent in-language sources correlated with mention rate far more strongly than page volume.

How many prompts do I need per language to trust the number?

Ten to fifteen high-intent prompts detect a real gap; forty tracks movement reliably. The binding constraint is repeat runs — five per prompt per engine minimum, because single-run readings vary by several points for reasons unrelated to your content.

Which engines should I track in each market?

At minimum one Google surface and one pure chat assistant, because they degrade differently: AI Overviews retained 63% of its English mention rate in German, ChatGPT 37%. Tracking only Google flatters your numbers; tracking only ChatGPT overstates the crisis. In Japan and Korea, add the locally dominant answer surfaces.

Which language gaps should I fix first?

Rank by revenue exposure, then by diagnosis speed. Category-term mismatch is cheapest (one to three weeks, largely a prompt-set and page-title correction) and worth resolving before commissioning content, because it changes what content you'd commission.

Does fixing German visibility hurt English visibility?

No evidence of it. Brand C's English rate held at 46–47% across a nine-week German programme, and no tracked brand showed an inverse relationship between language pairs. The corpora are largely independent — which also means gains don't transfer.

How long before multilingual AI search visibility actually moves?

Four to sixteen weeks depending on failure mode. Entity fragmentation and category-term mismatch move fastest; corpus absence and translation-shell content require earning third-party mentions and take a quarter. Watch cited-source coverage as the early signal.

Do non-Latin scripts need different handling?

Yes. In Japanese, katakana and Latin-script renderings of the same brand often behave as two weak entities rather than one strong one. Pick one canonical in-language name, use it identically everywhere, and verify by asking the same question with each spelling.


Written by

Founder of MaxAEO. Helping brands get found in AI search across ChatGPT, Perplexity, Google AI Overviews, and more.

Run a free AI visibility audit →