Blog

Could a Competitor With Only Public Data Have Written This Page?

Every content quality gate scores the finished text. None of them score the source. One question, asked before writing, catches what the others cannot.

By Anthony Gale — Co-Founder, Geoffy

AI DiscoveryContent StrategyGEO
Could a Competitor With Only Public Data Have Written This Page? cover image

Every quality gate in content marketing runs at the end. You write the page, then something scores it: a readability tool, a similarity check, an AI-detection pass, an editor.

All of those score the text. None score where the text came from. So a page can be well written, original by every similarity measure, technically clean, and still be a page four competitors will publish this quarter, because all five of you worked from the same sources.

There is a cheaper gate, and it runs before anything is written.

Could a competitor with access only to public sources have produced this page?

That is the whole test. One question, applied to the brief.

Why the source matters more than the text

Brand-owned pages are the majority of what AI assistants cite. Profound analysed 11.84 billion citations across eight models between April and July 2026, covering 3.02 million domains, and found roughly 57% of citations go to company-operated web properties. Social and UGC platforms sit far lower, and vary a lot by engine: Google AI Overviews cites social at 15.3%, while Microsoft Copilot uses a social source about once in every 29 citations.

The intuitive version of this story, that AI prefers Reddit and reviews to brand sites, is not what the largest dataset says.

That is good news and a problem at once. Good, because your own pages are the surface that gets cited. A problem, because so is everyone else’s.

Four inputs sit behind most ecommerce content briefs: the existing product page, Search Console, competitor pages, and general category knowledge. Every one is available to every competitor, and to every AI writing tool any of them points at their catalogue.

That is not a quality problem. You can write beautifully from public inputs. It is a structural one. The same source material produces convergent pages, and a convergent page gives an assistant no reason to pick you over four others saying the same thing.

Running the test

Take the brief before it becomes a page and answer four questions.

Source. Name the record this page draws on. If the honest answer is “the category” or “what competitors cover”, the source is public.

Counterfactual. Could a competitor working only from public sources produce this page? If yes, it will not differentiate you.

Specificity. Does the page carry a fact, a figure or a sentence that exists only in your records?

Destination. Is this at the right level, product, collection or site-wide? A page pitched at the wrong level fails even with a good source behind it.

Question two is the one that does the work. The others explain the answer.

Four worked examples

A collection page for “waterproof walking boots”. Source: category knowledge and a look at what three competitors did. Counterfactual: any of them could write it, and two already have. It fails. Re-brief it against the questions your own customers ask about waterproofing, drawn from your search log, support queue and returns reasons, and it becomes a page only you could write.

A product FAQ block generated from the specification sheet. Source: your own product data, which sounds private and mostly is not. Your spec sheet is on your public page and probably your distributors’ too. It fails, but only just. What rescues it is the questions: if the FAQ answers what buyers actually asked rather than what the spec implies, the source changes and the page passes.

A buying guide comparing your range. Source: your catalogue, all public. It fails. It passes once it carries the thing no competitor holds: which products get returned against each other and why, or which comparison your support team makes every week.

A page built from your zero-result search queries. Source: buyers typing into your search box and finding nothing. Not public, not purchasable, not inferable. It passes at the first question.

The pattern is consistent. Public inputs produce pages that are fine and interchangeable. First-party records produce pages that are hard to reproduce.

What to do with a page that fails

Do not delete it.

A page that fails the counterfactual is usually still worth having. Category pages need to exist and buying guides get used. The test does not sort content into keep and bin. It sorts content into cite-worthy and merely present. A page that fails is re-briefed against a first-party source, not thrown away. Often that means keeping the structure and replacing the evidence: same buying guide, comparison points now drawn from your returns log rather than a competitor’s table.

If a page genuinely has no first-party source available, publish it anyway and hold your expectations at the right level. It will do a job. It will not be the page that gets you named.

The evidence on generic content, stated honestly

There is a version of this argument that overreaches, and it is worth marking the line.

Ahrefs studied around 331,000 pages across 100,000 SERPs in July 2026 and found pages with high AI-generated content signals received two to three times fewer impressions and about nine percentage points lower indexation. A real finding, pointing the right way.

It is also correlational, and it relies on a proprietary detector. It does not show that generic content performs worse than publishing nothing. No study shows that, in AI retrieval or in search. And it is not evidence that Google penalises AI-generated content. The helpful content system was folded into core ranking in March 2024, and Google’s spam policy targets intent to manipulate rankings rather than the use of AI.

The honest claim is narrower and still strong enough: generic content is out-competed and under-indexed. It does not get you punished. It gets you passed over.

Where to use it

The test runs in about a minute per brief. Put it where a page gets commissioned, not where it gets reviewed, because by review time you have paid for the writing.

If you run content through an agency it also gives you a clean instruction that does not require reviewing drafts line by line: name the first-party record behind each brief. Briefs that cannot name one get re-scoped before anyone writes.

For how to choose between the tools that claim to measure any of this, see how to choose an AI visibility tool. For where the first-party records actually live, see Search Console cannot see your AI discovery problem and your returns data is a content brief.

Frequently asked questions

Is this the same as checking for duplicate content?

No. Duplicate checks compare finished text against text that already exists. This test asks whether the source material was available to anyone else, which catches pages that are entirely original in wording and entirely predictable in substance.

What counts as a public source?

Anything a competitor could reach without your permission: your product pages, your competitors' pages, published reviews, category knowledge, keyword tools, and Search Console data for their own property. Search Console feels private because it is your account, but the demand it describes is visible to everyone through keyword tools.

We use AI to write our content. Does this rule that out?

Not at all. The test is about the input, not the tool. An AI drafting a page from your returns log and your support queue is working from material no competitor holds. The same tool working from a competitor's page is not.

How does this relate to coherence?

Coherence tests the output — whether the visible page, the structured data and the rest of your catalogue agree with each other. Provenance tests the input. A page can be perfectly coherent and completely interchangeable, so both gates are worth running.

Next step

Ready to apply this to your catalogue?

Move from theory to implementation with parity-first GEO workflows.

Turn your catalogue into AI-readable product pages and structured data.

Already have an account? Sign in

Structured outputs enabled
First-party pages published
Discovery coverage expanding