Does Schema Markup Help AI Citations? What the Evidence Shows
Ahrefs found AI-cited pages carry JSON-LD almost 3x more often, yet its controlled test showed no causal uplift. What schema really does for AI visibility.
Structured data is machine-readable markup — schema.org vocabulary, usually written as JSON-LD — that describes what a page contains. The evidence that it drives AI visibility is mixed: Microsoft confirms its LLMs use schema, Google states no special markup is required for AI features, and Ahrefs found AI-cited pages carry JSON-LD almost 3x more often — yet its controlled experiment found no causal uplift. Schema helps AI systems understand your site; on its own, it does not make them cite you.
What is structured data, and how much of the web already uses it?
Structured data is code that labels the meaning of page content using the shared schema.org vocabulary, most often embedded as a JSON-LD block. It tells machines that this is a product, this is its price, this is the organization behind the site — without forcing them to parse prose.
It is no longer a niche practice. According to Web Data Commons at the University of Mannheim (December 2024), 51.25% of the 2.4 billion HTML pages in the October 2024 Common Crawl contain structured data, up from 5.7% in 2010. Some 11.5 million websites use JSON-LD — 70% of all sites that annotate — and sites with Product markup grew from 581,000 to 3.3 million.
The Web Almanac 2024 by HTTP Archive (November 2024) measured JSON-LD on 41% of mobile pages, up from 34% in 2022. The most common JSON-LD types are WebSite (12.73%), Organization (7.16%), BreadcrumbList (5.66%) and LocalBusiness (3.97%). Half the web already speaks this language. The real question is whether AI engines listen.
Which AI companies actually confirm they use schema markup?
Only Microsoft has confirmed that schema markup feeds the LLMs behind its assistant. Google’s public confirmations cover Search in general — for AI Overviews and AI Mode its documentation states that no special schema is required. OpenAI, Perplexity and Anthropic have confirmed nothing. At SMX Munich in March 2025, Fabrice Canel, Principal Product Manager at Microsoft Bing, confirmed that schema markup helps Microsoft’s LLMs — the models behind Copilot — understand web content, and recommended pushing fresh content through IndexNow, as reported by Search Engine Land (March 2025):
“Gen AIs value fresh content in particular, partly as a reference check of their LLM training data.” — Fabrice Canel, Microsoft Bing, at SMX Munich, March 2025
Google is more guarded. Its documentation on AI features (consulted July 2026) states that there is “no special schema.org structured data that you need to add” to appear in AI Overviews or AI Mode, and no additional machine-readable files either. The companion AI optimization guide (updated July 10, 2026) repeats that structured data is not required for generative AI search, while recommending you keep it for rich results as part of your general SEO strategy.
John Mueller of Google, asked whether extensive schema helps LLMs, answered “yes, no, and it depends”, according to Search Engine Roundtable (January 2026). Some features depend heavily on structured data — Shopping needs price, shipping and availability — while in others it only enriches the result. He noted this was not official guidance.
And the rest? As of March 2026, OpenAI, Perplexity and Anthropic have published no official confirmation that they use on-page schema markup in their answers, per Search Engine Land (March 2026).
Does adding schema markup increase AI citations?
The honest answer: the correlation is strong, but causation is unproven. Three datasets tell the story.
Ahrefs (May 2026) analyzed 6 million URLs and found that AI-cited pages were almost three times more likely to have JSON-LD than non-cited pages. The same team then ran a controlled experiment: 1,885 pages that added JSON-LD between August 2025 and March 2026, against a control group of 4,000 pages. Adding schema produced no major uplift in citations on any platform — AI Overviews -4.6%, AI Mode +2.4% (not significant), ChatGPT +2.2% (not significant).
A study by AirOps compiled by Analyzify (16,851 ChatGPT queries, 353,799 pages) leans the other way: pages with JSON-LD had a 38.5% citation rate versus 32.0% without it, a 6.5-point gap. The types with the highest citation rates were BreadcrumbList (46.2%), FAQPage (45.6%) and Organization (44.3%). Yet the same roundup includes contradictory results: OtterlyAI found no impact on Perplexity.
| Study | Sample | Finding | Limitation |
|---|---|---|---|
| Ahrefs correlation (May 2026) | 6 million URLs | AI-cited pages almost 3x more likely to have JSON-LD | Correlation only, no causation |
| Ahrefs controlled experiment (Aug 2025 – Mar 2026) | 1,885 test pages vs 4,000 control | No significant citation uplift on any platform | Single experiment, one time window |
| AirOps via Analyzify | 16,851 ChatGPT queries, 353,799 pages | 38.5% citation rate with JSON-LD vs 32.0% without | Correlational; contradicted by OtterlyAI on Perplexity |
In UpgradePro’s own labeling system: the 3x gap is MEASURED, but the claim that schema causes citations remains ASSUMED — and the one controlled test available failed to confirm it.
How does your markup reach an AI answer: live fetch or search index?
There are two separate paths, and schema only travels one of them. When a chatbot fetches your page live to compose an answer, it reads the visible HTML. When it leans on a search index — Google’s or Bing’s — your markup has already shaped how that index understands the page.
searchVIU (December 2025) ran 8 tests in October 2025 with a page whose fictitious prices were split between visible HTML, JavaScript and JSON-LD. None of the five systems tested — ChatGPT, Claude, Perplexity, Gemini and Google AI Mode — read the price that existed only in JSON-LD during a live fetch. Only Gemini rendered JavaScript.
“JSON-LD Schema is NOT read by AI chatbots during direct fetch” — searchVIU, December 2025
The practical implication: schema influences AI answers indirectly, through the search indexes that assistants query, not through the on-demand crawl. This is the nuance most vendor content skips. How assistants select and rank their sources is a topic of its own — we break it down in how AI engines choose what to cite.
Which schema types should a small business prioritize in 2026?
Prioritize Organization, Product if you sell, and Article with BreadcrumbList on content — and stop expecting anything from FAQPage in Google.
| Schema type | Best for | Evidence | Not recommended for |
|---|---|---|---|
| Organization | Consolidating your entity: name, url, logo, sameAs | 44.3% ChatGPT citation rate (AirOps); second most common JSON-LD type at 7.16% (Web Almanac 2024) | Expecting visible rich results from it alone |
| Product | E-commerce; structured price and availability | Shopping features depend heavily on structured data (Mueller, reported); ChatGPT shopping requires product feeds | Sites with no transactional pages |
| Article + BreadcrumbList | Blogs and content sites | BreadcrumbList had the top ChatGPT citation rate in AirOps (46.2%) | — |
| FAQPage | Q&A content, at low cost | 45.6% ChatGPT citation rate (AirOps) | Google rich results — gone entirely since May 7, 2026 |
Product is where structure has real teeth. OpenAI’s product feed spec for ChatGPT shopping (Agentic Commerce) requires a structured feed with mandatory fields — item_id, title, description, brand, price in ISO 4217, url, image_url, availability — which controls whether a product appears in ChatGPT searches and whether it can be bought via direct checkout. That is direct proof that AI commerce runs on structured data, even if through a feed rather than on-page JSON-LD.
FAQPage is a cautionary tale. Google restricted FAQ rich results to authorized government and health sites in August 2023, and since May 7, 2026 they stopped appearing entirely, with reports and the API retiring between June and August 2026, per Search Engine Journal (May 2026). Keep the markup if you have it; do not build a strategy on it.
All of this only pays off on a sound base: valid markup, served in the HTML, on crawlable pages. That base is what we cover in technical SEO foundations, and it is the first thing a technical SEO audit should check.
What role do entity SEO and knowledge graphs play?
Entity SEO shifts the unit of optimization from pages to entities: the companies, products and people your content is about. Consistent Organization markup — a stable @id, plus sameAs links to your real profiles — gives knowledge graphs an unambiguous identity to anchor, so systems can resolve who you are across the web.
The clearest signal that schema.org is becoming AI infrastructure is NLWeb. Microsoft announced it on May 19, 2025: an open-source project that turns websites into conversational AI interfaces, built on the schema.org and RSS data websites already publish. Its creator is R.V. Guha — who also created schema.org, RDF and RSS — and every NLWeb instance doubles as an MCP server, ready to talk to AI agents.
A site with clean markup is halfway to being agent-readable. Other machine-readable proposals, like llms.txt, aim at the same goal from a different angle — we compare adoption and evidence in our llms.txt guide.
What should you not expect from schema markup?
Do not expect schema alone to earn you citations. As of March 2026 there are no peer-reviewed studies on schema markup and AI visibility, per Search Engine Land (March 2026) — the available evidence comes from vendor and SEO-industry research, with the contradictions shown above.
Watch out for a common misattribution. The academic paper that coined GEO, Aggarwal et al. (KDD 2024), measured that tactics like adding source citations, statistics and quotations boost visibility by up to 40% in generative engine responses — visibility measured as Position-Adjusted Word Count, not as traffic. The paper did not test schema.org as a method. Blogs regularly attach that 40% to markup; the paper does not support that.
And remember Google’s own position: structured data is not required for generative AI search. Schema is understanding infrastructure, not a citation lever.
What should you do this week?
A realistic checklist for a small business, in order:
- Inventory and validate your existing markup. Broken or contradictory JSON-LD helps nobody; fix what is already there before adding more.
- Serve JSON-LD in the initial HTML, not injected by JavaScript. In searchVIU’s tests, only Gemini rendered JavaScript; anything JS-only is invisible to the rest during live fetches.
- Add or complete Organization markup on your homepage — name, url, logo, a stable @id and sameAs links to your real profiles — so knowledge graphs can resolve your entity.
- If you sell online, prepare structured product data. ChatGPT shopping requires a compliant feed, and Shopping features depend heavily on structured price, shipping and availability, according to John Mueller (Search Engine Roundtable, January 2026; not official guidance).
- Stop waiting for FAQPage rich results. They are gone from Google since May 7, 2026. Keep answer-shaped content visible on the page instead.
- Do not block the crawlers that feed the indexes. Schema reaches AI answers through the Google and Bing indexes; if they cannot crawl you, markup is moot.
- Measure before and after. Baseline your AI mentions before touching markup — otherwise you cannot tell signal from noise.
That last step is the sequence we follow in our AI visibility service: baseline audit first, implementation second, measurement against the baseline third. Schema markup earns its place in that plan — as infrastructure, not as a magic lever.
Frequently asked questions
Do AI engines read the JSON-LD on my website?
It depends on the path. In tests run in October 2025 and published in December 2025, searchVIU found that ChatGPT, Claude, Perplexity, Gemini and Google AI Mode all ignored JSON-LD during live page fetches and read only the visible HTML. Through the search index the story changes: Microsoft has confirmed that schema markup helps the LLMs behind Copilot understand web content. Google's public position is narrower: its documentation states that no special schema.org structured data is required for AI Overviews or AI Mode, and its publicly confirmed benefit is for Search in general.
Does Google require special schema markup for AI Overviews or AI Mode?
No. Google's official documentation states that there is no special schema.org structured data to add and no additional machine-readable files or requirements beyond standard SEO practices. Google still recommends keeping structured data as part of your overall SEO strategy because it continues to power rich results.
Is it proven that adding schema increases AI citations?
Correlation is strong, causation is not proven. Ahrefs found that AI-cited pages are almost three times more likely to carry JSON-LD, and AirOps measured a 38.5 percent citation rate with JSON-LD versus 32.0 percent without it, across 16,851 ChatGPT queries. But the controlled Ahrefs experiment on 1,885 pages found no significant uplift after adding schema. The honest reading: schema accompanies well-maintained sites, and on its own it is not a magic lever.
Which schema types should a small business prioritize?
Organization with sameAs links to consolidate your identity in knowledge graphs, Product data if you sell online since ChatGPT shopping runs on mandatory structured product feeds, and Article plus BreadcrumbList on content pages. BreadcrumbList, FAQPage and Organization showed the highest citation rates in the AirOps study.
Is FAQPage schema still worth adding?
Google no longer shows anything for it. FAQ rich results were restricted to authorized government and health sites in August 2023 and stopped appearing in Google Search entirely on May 7, 2026. The markup does no harm, and AirOps correlates it with ChatGPT citations, but do not expect rich results from Google for it.
What do structured data have to do with the agentic web?
A lot. NLWeb, the open-source Microsoft project created by R.V. Guha, the creator of schema.org, turns websites into conversational interfaces using the schema.org and RSS data a site already publishes, and every NLWeb instance also works as an MCP server for AI agents. A site with clean markup already has half the work done.