Adding schema markup will not, by itself, get your pages cited by ChatGPT, Gemini or AI Overviews. That is the finding of the first serious attempt to test the claim: a cross-platform study published in February 2026 collected 730 AI citations across 75 commercial queries and found schema present on 43.1% of cited pages and 44.8% of pages that were never cited. The markup simply did not separate the winners from everyone else. This matters because "add schema for AI" has become the most repeated piece of GEO advice on the internet, and a lot of budget is moving on the strength of it. What follows is what the evidence actually shows, what structured data genuinely does for you, which types are still worth the effort, and what predicts citation instead.
TL;DR
Generic schema markup does not predict whether an AI engine cites you. Markup that carries real attributes, such as Product and Review with populated prices, ratings and specifications, does correlate with citation, and rank position predicts it far better than either. Keep doing schema for the things it genuinely delivers, and stop buying it as an AI shortcut.
- The finding: schema appeared on 43.1% of AI-cited pages and 44.8% of uncited ones
- The exception: attribute-rich Product and Review markup was cited at 61.7% against 41.6% for generic
- The real driver: organic rank position, with each place down cutting citation odds by about a quarter
- Still worth doing: rich results, entity identity, authorship and dates, unambiguous facts
By the numbers
43.1%
of AI-cited pages carried schema, against 44.8% of uncited pages, across 730 citations from ChatGPT and Gemini. SSRN study
61.7%
citation rate for attribute-rich Product and Review markup with populated pricing, ratings and specs, against 41.6% for generic markup. SSRN study
15%
of all sources listed in Google AI summaries came from just three sites: Wikipedia, YouTube and Reddit. Pew Research
Industry figures are cited for context; outcomes vary by business and implementation.
What the evidence actually says
The study worth knowing about ran 75 commercial queries across five categories through ChatGPT with browsing and Gemini with search grounding, captured the 730 pages those engines cited, and compared them against the Google top ten for the same queries as a control, 1,006 unique pages in total. Every page was checked for JSON-LD schema and for domain authority. The first pass suggested schema was actively associated with fewer citations, which sounded dramatic; once the control set was corrected and errors clustered by query, the effect collapsed to nothing at all. That is the more useful result. Schema presence is so close to universal on commercial pages that it has almost no discriminating power. Everyone has it, so it cannot be what separates you.
One study is one study, and this one is a preprint testing correlation rather than a controlled experiment. But it lines up with what live testing keeps showing: when an answer engine fetches your page in real time, it largely reads the visible text. It is not hunting through your JSON-LD for a summary; it is lifting the sentence that answers the question.
Why attribute-rich markup is the exception
The one place structured data did move the numbers is instructive. Product and Review markup carrying populated pricing, ratings and specifications was cited on 61.7% of pages, against 41.6% for generic implementations. The likely reason is not that engines love schema; it is that those fields carry facts the prose often leaves out or buries. A price, a rating out of five, a release date and a set of specs are exactly the atoms an answer engine needs to build a comparison, and a page that states them unambiguously is a cheaper source than one where the model has to infer them. Empty markup that only says "this is an Article" adds no facts, so it adds no advantage.
The practical translation: schema earns its keep when it publishes data, not when it labels a template. If your Product markup has no price in it, you have written a comment, not a signal.
So why keep doing schema at all?
Because it still does four jobs nothing else does, and none of them are AI citations.
- Rich results: star ratings, prices, FAQs where still supported, events, recipes and breadcrumbs come from markup, and they change how your listing looks in a results page that is already showing fewer of them
- Entity identity: Organization, Person and sameAs links tell every system that this business, this author and these profiles are the same entity, which is how you get treated as a known brand rather than an unknown URL
- Provenance: author and date fields tie content to real people and verifiable timestamps, which is increasingly what separates a source worth quoting from an anonymous page
- Clean extraction: prices, availability, opening hours and locations get read the same way by every parser, so nothing depends on a model guessing correctly from your prose
Those are strong reasons. They are just not the reason it usually gets sold. When we do technical SEO work on a site, schema is a hygiene layer we expect to find working, in the same category as canonical tags and a valid sitemap, not a growth lever we bill as a strategy.
What predicts citation instead
Rank first. In the same study, organic position was the dominant predictor: pages at position 1 were cited in roughly 43% of queries against 5% at position 7, with each place down the page cutting the odds by about a quarter. Answer engines are, to a large extent, reading the same results you are. Then comes recognisable authority, and the shape of that is visible in Pew's data on Google AI summaries, where Wikipedia, YouTube and Reddit alone supplied 15% of all listed sources. Engines lean on entities they already trust. Third, extractability: a page that answers the question in its first two sentences, in plain declarative prose, gives a model something clean to quote, while a page that warms up for six paragraphs does not. Fourth, third-party mentions, because a claim repeated on sites the engine trusts is a claim it will repeat too.
That ordering is why GEO work that starts with markup usually disappoints, and why the same work starting with rankings, answer-first rewriting and earned mentions tends not to. There is more on that sequencing in our guide to GEO and AEO, and on the traffic maths behind it in what still works under zero-click search.
A sensible schema plan for this quarter
Spend an afternoon, not a retainer. Validate what you already have and fix the errors, because broken markup is worse than none. Fill in the attributes on the types that carry facts, especially Product, Offer, Review, LocalBusiness and Event, and delete markup that declares a type and nothing else. Get Organization and Person right once, site-wide, with accurate sameAs profiles. Make sure every article has a real author and a real date. Then stop, and put the rest of the time into the things that actually move citation: better rankings on the queries you care about, answers written to be lifted, and being mentioned somewhere other than your own domain.
Bottom line: schema markup is worth doing and is not a shortcut to AI citations, and the evidence now says so out loud. Publish real facts in your markup, get the fundamentals right, and treat any pitch that sells structured data as a generative-search growth strategy with the scepticism it has earned.
Frequently asked questions
Does schema markup help you get cited by AI?
Not on its own. A cross-platform study of 730 AI citations from ChatGPT and Gemini found schema was present on 43.1% of cited pages and 44.8% of pages that were not cited, a difference small enough to be statistically indistinguishable. Markup that carries real attributes, such as Product and Review schema with populated prices, ratings and specifications, did show an effect, but empty boilerplate markup did not.
Which schema types matter most for AI search?
The types that publish facts a machine cannot infer from prose: Product and Offer with real prices and availability, Review and AggregateRating with genuine ratings, Organization and Person for entity identity, Article with verifiable authors and dates, and Event, Recipe or LocalBusiness where they apply. In the study, attribute-rich Product and Review markup was cited on 61.7% of pages against 41.6% for generic implementations.
Should I still add schema markup in 2026?
Yes, but for the right reasons. Schema still earns rich results in Google, confirms entity identity across the web, ties content to real authors and dates, and removes ambiguity for any system parsing your page. Those are solid reasons. Expecting markup alone to buy AI citations is not, and any agency selling schema as a GEO shortcut is overselling it.
What actually predicts whether AI engines cite a page?
Google organic rank position is the strongest single predictor in the available research: pages at position 1 were cited in about 43% of queries against 5% at position 7, with each place down the results reducing the odds by roughly a quarter. After rank come recognisable authority, clear extractable prose that answers the question directly, and third-party mentions. Structured data supports those signals rather than substituting for them.