How AI Engines Cite Sources: ChatGPT, Perplexity, Gemini, Claude
Which indexes ChatGPT, Perplexity, Google AI Overviews and Claude cite from, the most-cited domains, and the changes that raise visibility up to 40%.
AI engines cite sources in two steps: index retrieval, then passage selection. The index differs per engine — Google’s own index for AI Overviews and Gemini, Perplexity’s proprietary index of more than 200 billion URLs, third-party search providers plus OAI-SearchBot for ChatGPT, and, the public evidence suggests, Brave Search for Claude. Selection then favors pages with quotable statistics, direct quotations and clear sourcing. Google rankings alone are not enough: only 12% of AI-cited URLs rank in Google’s top 10 for the same prompt.
Why does it matter how AI engines choose sources?
Citations are the only channel through which AI answers send readers, and attribution, back to your site. Attribution is valuable enough that OpenAI pays for it: according to OpenAI (March 2024), its content deals with publishers such as AP, Axel Springer, Financial Times, GEDI (La Repubblica, La Stampa), Le Monde, News Corp, Prisa Media (El País), Reuters, Time and Vox Media put their content in ChatGPT as summaries “with attribution and enhanced links to the original articles”.
The clicks that do arrive may also behave differently. According to Google Search Central, clicks coming from results pages with AI Overviews are higher quality, meaning users tend to spend more time on the site. That is Google’s own claim about its own feature — but it matches why citation share is now a metric worth tracking alongside rankings.
Which search index does each AI engine use?
Four engines, three different retrieval stacks, and one educated guess. Knowing which index feeds each engine tells you exactly which crawler must be able to reach your pages.
ChatGPT: third-party search plus its own crawler
ChatGPT search launched on October 31, 2024. According to OpenAI’s announcement, it leverages third-party search providers — historically Bing — as well as content provided directly by partner publishers. OpenAI’s current page names no provider, so treat Bing as the historical origin, not a confirmed present-day fact.
On top of that, OpenAI’s bot documentation separates three crawlers with distinct jobs: OAI-SearchBot powers search in ChatGPT and decides which sites surface in results; ChatGPT-User performs visits a user triggers; GPTBot collects content for model training. A site that blocks OAI-SearchBot in robots.txt does not appear in ChatGPT search answers.
Perplexity: its own index, built to cite fragments
Perplexity operates its own crawler, PerplexityBot, which per the official documentation exists only to surface and link websites in results — it is not used to crawl content for foundation-model training. A separate agent, Perplexity-User, handles user-initiated visits and may ignore robots.txt.
The scale is unusual for a startup: according to a Perplexity Research article (July 15, 2026), its search infrastructure tracks more than 200 billion unique URLs, combines lexical and semantic retrieval, and treats document sections and fragments as first-class units so the engine can cite the most atomic passage possible. The practical consequence: self-contained sections that answer one question each are what this retrieval design rewards.
Gemini and AI Overviews: Google’s index plus query fan-out
Google is unusually explicit about the entry requirements.
“There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” — Google Search Central, AI Features and Your Website (consulted July 2026)
Being indexed and eligible for a snippet is the whole ticket. AI Overviews and AI Mode then apply a technique Google calls query fan-out: issuing multiple related searches across subtopics and data sources, which according to Google surfaces a broader and more diverse set of links than classic search. One more distinction matters: Google-Extended controls whether your content trains or grounds Gemini — it does not control appearance in Search’s AI features.
Claude: the Brave Search trail
Anthropic has never named its search provider. The evidence points at Brave Search: TechCrunch reported (March 21, 2025) that Anthropic added Brave Search to its subprocessor list and that the feature’s code contained a BraveSearchParams parameter; TechCrunch also found at least one search where Claude and Brave returned identical citations. Anthropic itself runs three separately blockable crawlers, per its Help Center: ClaudeBot (training), Claude-User (user-initiated visits) and Claude-SearchBot (search result quality).
Which domains do AI engines cite most often?
Wikipedia appears in the top list of all three major platforms and is the number-one domain on ChatGPT (16.3%), though Ahrefs measures YouTube above it on Perplexity and AI Overviews. Each engine keeps a distinct source diet. Ahrefs (June 11, 2025) analyzed 78.6 million interactions with June 2025 data:
| Domain | ChatGPT | Perplexity | Google AI Overviews |
|---|---|---|---|
| Wikipedia | 16.3% | 12.5% | 8.4% |
| YouTube | — | 16.1% | 9.5% |
| — | — | 7.4% | |
| Quora | — | — | 3.6% |
| Reuters | 4% | — | — |
Share of mentions in the Ahrefs study; ”—” means the domain did not stand out in that platform’s top list. ChatGPT’s tilt toward news media such as Reuters is consistent with OpenAI’s publisher agreements; AI Overviews is the only platform of the three where forums — Reddit and Quora — enter the top.
Profound (June 2025, updated August 2025) measured 680 million citations from August 2024 to June 2025 and found the same shape at different magnitudes: Wikipedia is ChatGPT’s number-one source at 7.8% of all citations, Reddit leads on Perplexity (6.6%) and on Google AI Overviews (2.2%), .com domains concentrate 80.41% of citations and .org domains 11.29%.
Why do Ahrefs and Profound disagree on Wikipedia’s share for ChatGPT (16.3% vs 7.8%)? Different methodologies, metrics and time windows. Neither number is wrong; each is only valid inside its own study. Any vendor quoting a single “share of AI citations” without naming the study is selling you noise.
How stable are these citation patterns?
Not very — at least not on ChatGPT. Semrush (November 10, 2025) tracked more than 230,000 prompts and over 100 million citations between July 14 and October 12, 2025. Reddit fell from appearing in roughly 60% of ChatGPT answers to roughly 10% by mid-September 2025, and Wikipedia dropped from about 55% to under 20%. Google AI Mode stayed stable over the same window, with LinkedIn (~15%) in front.
The lesson for a small business: a one-month snapshot of “who gets cited” is not a strategy. Citation shares on ChatGPT can move by dozens of points in weeks, so any serious effort needs repeated measurement — a baseline, an intervention, and an after.
Does ranking in Google guarantee AI citations?
No. Ahrefs studied 15,000 long-tail queries (August 11, 2025) and found that only 12% of the URLs cited by AI assistants rank in Google’s top 10 for the original prompt.
“80% of those citations don’t rank anywhere in Google for the original query.” — Ahrefs, AI search overlap study (August 2025)
The overlap varies by engine: Perplexity is the most Google-aligned at 28.6%, while ChatGPT sits around 8%. So existing rankings transfer somewhat to Perplexity and barely to ChatGPT. This gap is the practical argument for treating GEO as a discipline distinct from classic SEO: the two channels reward overlapping but different work, and one does not automatically buy the other.
Which content changes actually increase citation probability?
The strongest controlled evidence remains the academic GEO paper (Aggarwal et al., KDD 2024), tested on GEO-bench, a benchmark of 10,000 queries. It showed that optimizing content can raise visibility in generative engines by up to 40%. One nuance the summaries usually drop: that 40% is measured on Position-Adjusted Word Count — how much of the answer cites you and how prominently — not on clicks or traffic.
According to the full text, the most effective tactics were:
- Adding direct quotations from sources: ~27.8% improvement.
- Adding statistics: ~25.9%.
- Citing sources: ~24.9%.
- Improving fluency of the writing: ~25.1%.
Keyword stuffing — the reflex of old SEO — did not work. And the distribution of gains is the part small businesses should read twice: the lowest-ranked sites gained the most, up to +115% visibility. Generative answers redistribute attention downward, which is precisely why it is worth learning what makes content citable before your bigger competitors do.
Should you allow or block AI crawlers in robots.txt?
Separate search bots from training bots — they are different user agents with different consequences.
| Company | Search / citations | User-initiated visits | Training |
|---|---|---|---|
| OpenAI | OAI-SearchBot | ChatGPT-User | GPTBot |
| Perplexity | PerplexityBot | Perplexity-User (may ignore robots.txt) | — (not used for foundation models, per docs) |
| Anthropic | Claude-SearchBot | Claude-User | ClaudeBot |
| Googlebot (index + AI features) | — | Google-Extended (Gemini training/grounding) |
Blocking OAI-SearchBot removes you from ChatGPT search answers; blocking GPTBot or ClaudeBot only opts you out of training, without touching your presence in cited answers. On Google’s side, the classic controls — noindex, nosnippet, data-nosnippet, max-snippet — also limit your appearance in AI features, while Google-Extended governs Gemini training and grounding only. User-agent strings and copy-paste robots.txt examples are in our AI crawlers guide.
What should an SMB do this week?
Five actions, all doable without buying any tool:
- Audit robots.txt. Confirm OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot are allowed. Decide on GPTBot and ClaudeBot separately — that is a training-consent decision, not a visibility one.
- Check snippet eligibility in Google. Hunt for unintended noindex, nosnippet or max-snippet directives; per Google, indexation plus snippet eligibility is the only entry requirement for AI Overviews.
- Upgrade your five most important pages with sourced statistics, direct quotations and explicit source citations — the three tactics with ~25-28% measured lift in the GEO paper.
- Make sections self-contained. Question-style headings with a direct answer in the first sentences, because Perplexity’s retrieval treats fragments as first-class citation units.
- Record a baseline. Run your 10-20 core buyer prompts on ChatGPT, Perplexity and Google monthly and log which domains get cited; Semrush showed those shares can shift dozens of points within weeks.
If you want that baseline measured and labeled honestly — what is measured, what is calculated, what is assumed — that is exactly what our AI visibility audit delivers.
Frequently asked questions
Which search index does each AI engine use?
ChatGPT combines third-party search providers (historically Bing), its own OAI-SearchBot crawler and content from partner publishers. Perplexity builds and operates its own index, tracking more than 200 billion unique URLs with PerplexityBot. Gemini and AI Overviews use Google's index. Claude almost certainly relies on Brave Search: Anthropic listed Brave as a subprocessor in March 2025, though it has never confirmed the relationship officially.
Do I need special markup or a special file to appear in Google AI Overviews?
No. Google states in its official documentation that there are no additional requirements and no special schema: your page needs to be indexed and eligible to show a snippet in Search. The classic controls — nosnippet, data-nosnippet, max-snippet and noindex — also limit how you appear in AI features.
If I already rank number one in Google, will AI engines cite me?
Not necessarily. According to an Ahrefs study of 15,000 long-tail queries, only 12% of the URLs cited by AI assistants rank in Google's top 10 for the same prompt, and 80% of the citations do not rank anywhere in Google for the original query. Perplexity is the most aligned with Google, at 28.6% overlap; ChatGPT sits around 8%.
Which websites do AI engines cite most?
Wikipedia, YouTube and Reddit dominate, but each platform has its own pattern: ChatGPT favors Wikipedia and news media (OpenAI has content deals with several publishers), Perplexity favors YouTube and Reddit, and Google AI Overviews gives more room to forums like Reddit and Quora. These shares are volatile: Semrush measured Reddit falling from roughly 60% to roughly 10% of ChatGPT answers within weeks.
What concrete content changes increase my probability of being cited by AI?
The academic GEO paper (KDD 2024) found that optimizing content can raise visibility in generative engines by up to 40%. Measured on Position-Adjusted Word Count — how much of the answer cites you and how prominently, not clicks or traffic — the most effective tactics were direct quotations from sources (~27.8% improvement), statistics (~25.9%), source citations (~24.9%) and better fluency (~25.1%). Traditional keyword stuffing did not work. Lower-ranked sites gained the most, up to +115% visibility.
Should I allow or block AI bots in my robots.txt?
Separate search from training. If you want to be cited, allow the search bots: OpenAI's OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot. Blocking OAI-SearchBot removes you from ChatGPT search answers. The training bots — GPTBot and ClaudeBot — can be blocked independently without losing visibility in cited answers.