Executive summary
Through 2024 and 2025, advice on AI visibility ran ahead of evidence. That gap closed materially in May 2026, when two controlled studies — one measuring what doesn’t move AI retrieval, one measuring what captures citations — were published within ten days of each other. Read together, they frame the discipline: structure alone has no measurable effect, and coherent brand-owned pages win the citation. This report assembles that evidence, the concentration data that raises the stakes, and the mechanism that explains both.
1. The two findings that frame 2026
The null result. In May 2026, Ahrefs published a controlled study of schema-only optimisation: 1,885 treated pages against 4,000 matched controls, evaluated with a difference-in-differences design. The measured effect on AI retrieval — +2.4% in Google’s AI Mode, +2.2% in ChatGPT — was, in the authors’ words, statistically indistinguishable from zero. Adding structured data markup to otherwise unchanged pages did not move whether AI systems retrieved or cited them. The authors were careful to carve out what schema still does: it feeds entity recognition and knowledge graphs — the upstream registration layer. What it does not do is substitute for the page itself being worth citing.
The capture result. The same month, BrightEdge measured where consideration-stage citations actually go. Across eight industries, brand-owned commercial pages captured between 42% and 79% of citations at the consideration stage of the buying journey. The popular assumption — that AI answers are built almost entirely from third-party editorial and forum content — does not survive contact with this data. When a brand’s own page carries consistent, structured, checkable information, engines cite it directly.
One study shows the shortcut doesn’t work. The other shows the prize for doing the real work. The variable separating them is not markup, budget, or domain authority. It is whether the brand’s information is coherent enough for a model to stake an answer on.
2. The mechanism: why contradiction reads as risk
An AI assistant answering “what’s the best magnesium for sleep?” has to compose a confident recommendation in a few hundred tokens, usually with no prior context. To do that it reads what it can find — the product page, structured data, an FAQ, a marketplace listing, a review page — and resolves everything into one description.
If those sources agree, the model can describe the product specifically and commit. If they disagree — one dosage on the product page, another in an old blog post; one positioning on the homepage, another in the feed — the model faces a choice it is built to avoid: assert something that might be wrong. So it hedges, or it omits, and recommends the competitor next door whose sources don’t disagree with themselves.
This is why the Ahrefs result should surprise nobody. Schema is a promise about the page. If the page, the feed, and the wider web don’t keep that promise consistently, the markup adds a claim without adding confidence. And it is why the BrightEdge result is the encouraging half: the model prefers an owned page it can trust, because a coherent primary source is the lowest-risk citation available.
Earlier academic work pointed the same way. The Princeton-led GEO study (ACM KDD 2024) found structured, explicitly attributed content increased visibility in generative responses by up to 40% against unstructured equivalents — structure with substance, measured together, not markup alone.
3. The concentration effect: no long tail in an answer
Coherence would matter less if AI answers were generous. They are not.
In benchmark testing of roughly 3,800 supplement-category prompts run against five engines between January and April 2026 (2026 Supplements AI Visibility Index, 5WPR), four brands captured an estimated 47% of all observed citations. In separate, narrower category testing published earlier (Avenue Z, September 2025, three engines), a substantial minority of tested brands appeared zero times — not ranked low; absent.
Search never behaved this way. The results page had ten links, then more below, then a second page; being twelfth was survivable. An answer has three to five names in it. The competition is winner-take-most, and it resolves faster than search’s slow rank decay — because there is nothing below the fold of a sentence.
Concentration also compounds the coherence effect in both directions. Brands the models can describe confidently get cited, which produces more consistent third-party coverage, which strengthens the next retrieval. Brands the models skip generate no such record. The gap widens without anyone deciding it should.
4. Fragmentation makes coherence the only portable strategy
A year ago, a plausible response was: learn one engine’s quirks and optimise for it. That assumption has expired. ChatGPT’s share of generative AI traffic fell from roughly 76% to roughly 53% in a year (Similarweb), while Gemini’s share more than doubled and Perplexity and Copilot grew into real volume. Discovery is not consolidating onto one engine; it is splintering across several, each retrieving differently and each changing frequently.
Per-engine tactics therefore don’t scale — four engines, each moving every few weeks, cannot each be gamed indefinitely. What transfers across all of them is the property none of them can ignore: information that agrees with itself wherever it’s read. Coherence is engine-independent. It pays off in every engine at once, including those that haven’t launched yet.
5. What coherent brands do differently
Across first-party visibility audits — including a sustained audit programme in the supplements category — the brands that perform in AI answers share observable habits rather than budgets:
One story, everywhere. The product page, the FAQ, the reviews vocabulary and the structured data use the same claims and the same language — often because a single voice wrote them and kept them matched.
Comparison-ready specificity. Attributes stated explicitly (dose, form, material, certification), in terms a model can check against a competitor, rather than dissolved into lifestyle copy.
Maintained, not launched. Prices, availability and claims that stay synchronised as the catalogue changes. Coherence decays by default; the winners treat it as infrastructure with upkeep, not a project with an end date.
A recurring, uncomfortable pattern: brands with the deepest content libraries — years of SEO investment, agency retainers, topic clusters — are routinely beaten in AI answers by smaller competitors whose entire site is a handful of crisp, repeating, machine-readable claims. Search rewarded depth. AI rewards agreement. They are different optimisation targets, and success at the first confers nothing at the second.
6. Implications
For brands, the order of work follows from the evidence. First, look: ask three engines the question your best customer asks and read the answers side by side — presence, accuracy, consistency. Second, reconcile: find where your own surfaces contradict each other, because that is what the model sees. Third, structure what’s true: schema and discovery files built from a coherent catalogue, not painted over an incoherent one. Fourth, maintain: schedule the checking, because drift is the default state of any live catalogue.
For the industry, the May 2026 studies should retire two comfortable narratives at once: that AI visibility is a markup checkbox, and that it is entirely at the mercy of third-party coverage. The evidence says the decisive surface is the one the brand already owns — held to a standard most catalogues were never built to meet.
Conclusion
The judge has changed. Search’s judge tolerated internal contradiction because it ranked pages against each other; the assistant’s judge cannot, because it has to say something true in two sentences. Structure gets you read. Coherence gets you chosen. The brands that internalise the difference earliest will spend the next two years being the answer — and the rest will spend them wondering where the demand went.
Geoffy builds coherence infrastructure for AI commerce: structured discovery on the merchant’s own domain, kept in sync as the catalogue changes. Live today on Shopify and WordPress / WooCommerce.
References
- Ahrefs, Schema markup and AI retrieval: a controlled study, 11 May 2026.
- BrightEdge, consideration-stage citation analysis across eight industries, 21 May 2026.
- 5WPR, 2026 Supplements AI Visibility Index (Jan–Apr 2026 testing).
- Avenue Z, supplement AI visibility testing, September 2025.
- Similarweb, generative AI traffic share data, 2025–2026.
- Aggarwal et al., GEO: Generative Engine Optimization, ACM KDD 2024.