Multilingual GEO: How AI Citations Change With the Language You Ask In
AI engines cite different sources in Italian, Spanish or English — and the click benchmarks most marketers quote were measured only in the US. Why native-language measurement is not optional.
Ask an AI engine the same question in English, Italian or Spanish and you often get different cited sources. Independent research shows models pivot to English pages when local sources are thin, and the click-through benchmarks most marketers quote were measured only in the United States. For a multilingual European business, that means every imported US number carries a language ceiling — and native-language measurement is not optional.
Why doesn’t the click-through benchmark everyone quotes apply to your market?
The single most-cited number in generative search is a US number. Pew Research (July 2025) tracked the browsing of 900 US adults in March 2025 and found that when a Google AI summary appeared, users clicked a traditional result on 8% of visits, versus 15% when no summary appeared; only 1% clicked a source link inside the summary itself. It is a precise, useful figure — for English-speaking America. It says nothing about Milan or Madrid, because Google only switched on AI Overviews in Italy and Spain on 26 March 2025, per Search Engine Land (March 2025) — the same month Pew’s data window closed. Many benchmarks now circulating literally could not have measured those markets. We treat US click-through and overlap figures as directional in our GEO vs SEO breakdown, never as your number.
Do AI engines really cite different sources depending on the query language?
The strongest independent evidence comes from academia, not vendors. A Stanford-led study (Suzgun et al., arXiv:2605.22785, February 2026) put six commercial chatbots — Gemini 3 Flash and Pro, Grok 4, Claude 4.5 Sonnet, GPT-5 and GPT-4o mini — through 2,100 same-day news questions across six regional BBC services, 12,600 model-question instances in all. The retrieval pattern was lopsided:
“Nine of ten are primarily English-language despite four of six regions covering non-English content” — Suzgun et al., Evaluating Commercial AI Chatbots as News Intermediaries, February 2026
Nine of the ten most-cited domains globally were English-language, even though four of the six regions covered non-English content. One caveat we keep front and centre: this paper measures factual accuracy on news, not brand or commercial citations. Use it as evidence of the mechanism — an English retrieval bias — not as a study of how businesses get cited.
Is this a comprehension problem or a retrieval problem?
Retrieval, and the distinction matters. The models understand the languages fine; they fail at finding good local sources to ground their answers. In the Stanford data, accuracy on Hindi-language questions fell to roughly 79%, against 88.9-91.3% for the other regions — not because the model misread Hindi, but because it reached for English pages when well-indexed local sources were missing.
“a retrieval-and-grounding failure in which models pivot to English-language sources” — Suzgun et al., February 2026
More than 70% of all errors in the study were retrieval failures. The practical lesson: if your language’s web is thinly indexed on a topic, the engine defaults to English — and to whoever is well-sourced in English. That retrieval pipeline, query fan-out and all, is the same one we unpack in how AI engines pick their citations.
What do the GEO vendors add — and how far can we trust it?
Two vendors have published the largest multilingual citation analyses so far. Treat their numbers as informed and directional, not reproducible: both are interested parties, and neither releases an open dataset.
Profound (April 2026) analysed 3.25 billion citations across seven models and 14 countries, filtering every prompt by the country’s native language, and reached a blunt conclusion:
“a country’s geography is secondary to a much more powerful force: the language of the query” — Profound, April 2026 (vendor data)
Temso AI (early 2026) went further on the source-language question, analysing 7,058,891 citations across four models, six non-English languages (Italian and Spanish among them) and 47 sectors. Its “local-language citation rate” — the share of cited pages whose language matches the prompt — splits the models wide open:
| AI surface | Local-language citation rate |
|---|---|
| Google AI Overview | 85.4% |
| Microsoft Copilot | 76.7% |
| ChatGPT | 70.2% |
| Grok | 51.7% |
Source: Temso AI, early 2026 (vendor data, no public dataset).
That is a 34-point gap between the most local model and the least. Temso reports the same split inside a single language: for Dutch prompts, Grok cited more English sources (53.5%) than Dutch (38.3%), while the identical prompts in Google AI Overview returned 81.2% Dutch. Sector matters too — local-language citation ran at 76.9% for K-12 education but just 35.5% for hotels and hospitality, because globalised verticals lean on English content. The honest gap: there is still no independent, non-vendor study measuring brand citation by language on an open dataset. The mechanism is well-evidenced; the commercial figures are vendor-supplied.
How many Europeans actually use generative AI — and does a small base mean no opportunity?
Enough to matter, and unevenly. Eurostat (December 2025) reports that 32.7% of people aged 16-74 in the EU used generative AI in the three months before the survey. The spread is enormous: Denmark (48.4%), Estonia (46.6%) and Malta (46.5%) lead, while Italy sits near the bottom at 19.9%, second-lowest after Romania (17.8%). A low base is not an absent market — it is an early one, with less entrenched competition for citations. We track those adoption curves in AI search adoption in Europe.
What does this mean for every US statistic in this blog?
Every one of them carries a language ceiling. A click-through rate, a citation-overlap percentage or a conversion figure measured on US English-language traffic tells you about a market that had AI Overviews for months longer than Europe did, in the language the models retrieve most abundantly. It is a legitimate directional signal. It is not your baseline. Where we cite a US-only figure, treat it as observed-in-the-US and assumed-elsewhere — an assumption, not a measurement, until someone measures your language. Extending any of this to a smaller market such as Slovenian, which none of these studies covered, is pure extrapolation and should be labelled as such.
What should a multilingual SME do about it?
A sequence that respects the evidence, ordered by leverage:
- Measure in the market’s real language. Ask ChatGPT, Perplexity and Google (with AI Overviews) your customers’ questions in Italian, Spanish or Slovenian — not in English — and log who gets cited. An English baseline predicts the wrong thing.
- Check the answer per model, not in aggregate. The same Dutch prompt swung from 38.3% to 81.2% local sources depending on the engine (Temso); assume your language behaves differently on Grok than on Google AI Overview.
- Publish native, indexable content. Not English translated on the fly, but pages written in the market’s language that a crawler can fetch and parse — so the model has a local source to ground on instead of pivoting to English.
- Down-weight imported benchmarks. Use US numbers to understand direction, then replace them with your own measured baseline.
The through-line is simple: because the language of the query changes what gets cited, AI visibility has to be measured in the real language of each target market, not extrapolated from English. That principle — audit first, measure in-language, label every number as measured, calculated or assumed — is exactly how our AI visibility service approaches it, on the same SEO foundations that let a crawler reach your pages in the first place.
Frequently asked questions
Do AI engines answer and cite different sources depending on the language you ask in?
The evidence points to yes. An independent academic study (Stanford et al., arXiv:2605.22785, February 2026) found that nine of the ten most-cited domains globally were English-language, even though four of the six regions it tested covered non-English content: when models cannot find well-indexed local sources, they pivot to English. Vendor analyses from Profound and Temso AI, with published methodology, point the same way — the language of the query changes what gets cited, and each model behaves differently. Those are vendor figures (not independently reproducible), but consistent with the academic source.
If I translate my site into English, will that be enough to appear in Italian or Spanish answers?
It is not equivalent. Temso AI (a vendor) reports that the 'local-language citation rate' reaches 85.4% in Google AI Overview: when the question is in Italian, most cited sources are also in Italian. Native content tends to perform better than translated English for non-English queries. We cannot, however, promise a specific improvement figure — it depends on the model, the sector and the market.
Do figures from US studies (clicks, citation overlap) apply to my market?
They carry a language ceiling. The reference study on clicks (Pew Research, July 2025) tracked only 900 US adults with browsing data from March 2025. Google switched on AI Overviews in Italy and Spain on 26 March 2025, so many benchmarks now circulating could not even have measured those markets. Use them as directional orientation, not as your number.
Which AI model is most 'local' in the sources it cites?
According to Temso AI (a vendor), Google AI Overview is the most oriented toward local-language sources (85.4%) and Grok relies most on English (51.7% local). The same Dutch-language prompt produced 81.2% local sources in Google AI Overview but only 38.3% in Grok. The choice of engine matters as much as the language.
How many people use generative AI in Italy, Spain and other European markets?
Per Eurostat (official, 2025), 32.7% of Europeans aged 16-74 used generative AI in the previous three months. Italy is among the lowest in the EU (19.9%), meaning a small base but a lot of headroom; the Nordic countries exceed 46%. This is real market context, not an estimate.
Does UpgradePro measure AI visibility in four languages?
The honest answer is not to promise a figure or a scope that is not confirmed. What is well-founded is the principle: because the language of the query changes the citations, any AI-visibility measurement has to be done in the real language of the target market, not extrapolated from English data.