Blog

Your Returns Data Is a List of the Facts Your Product Pages Left Out

A return is a buyer who decided with incomplete information. Read as a content brief rather than an ops problem, the returns log names the missing fact, product by product.

By Anthony Gale — Co-Founder, Geoffy

AI DiscoveryProduct DataShopifyGEO
Your Returns Data Is a List of the Facts Your Product Pages Left Out cover image

Returns get treated as an operations problem. Reduce the rate, speed up the refund, work out which suppliers cause the most trouble. All sensible, and all of it looks backwards at a transaction that already went wrong.

There is a second reading, and almost nobody uses it.

A return is a record of a buyer who made a decision with incomplete information. They read the page, they formed an expectation, the product did not match it. Something they needed to know was not on the page, or was there and did not register.

Which makes your returns log a list of the facts your product pages left out. Written by buyers. Product by product. In their words.

What the return reason actually tells you

Take the most common codes on a fashion or homeware store: size too large, size too small, not as described, wrong item, defective, colour, style, unwanted.

An ops team reads “size too small” as a sizing problem and sends it to merchandising. Fair enough. But read the same code as a content signal and it says something more useful: this page did not carry the measurement, the fit note or the comparison a buyer needed to get the size right first time. The information exists somewhere in your business. It did not make it onto the page.

“Not as described” is blunter still. That is a buyer telling you the page and the product disagreed. It is the single most direct statement of a content failure available anywhere in ecommerce, and it arrives already attached to a specific SKU.

Now ask what an AI assistant does with that same page.

An assistant answering “will this fit a standard 60cm cabinet” or “is this warm enough for winter walking” has to find the fact in the page it retrieved. If the fact is missing, the assistant does one of two things. It stays quiet about your product and names one that does carry the detail. Or it infers, and describes your product in terms you never wrote.

The gap that produced the return and the gap that keeps you out of the answer are the same gap. One costs you a refund and a restock. The other costs you a mention you never knew was available.

Be honest about how thin the native data is

This is where most articles would tell you your returns data is an untapped goldmine. It is not, and overselling it would ruin the argument.

On Shopify, the returns reason is a short closed list. The ReturnReason enum offers colour, defective, not as described, size too large, size too small, style, unwanted, wrong item, other and unknown. That is it. Free text is limited too: returnReasonNote caps at 255 characters and the customer’s own note at 300.

A closed list of ten codes and a couple of short text fields is not a rich corpus. Say so plainly, because a brand that opens its returns export expecting essays will close it again in five minutes.

Some returns platforms are more generous. Loop Returns exposes a free-text return_reason through its warehouse reporting endpoint, so brands running Loop have real sentences rather than codes. Worth knowing which side of that line you are on before you plan any work.

Thin still beats absent

Here is why the thin signal is worth reading anyway.

It names the product. Not a category, not a segment. This SKU, this failure, this buyer. Almost nothing else in your content workflow operates at that resolution.

It names the failure mode. “Size too small” on one product and “not as described” on another are different content jobs, and the code tells you which one you have before you open the page.

And it is a record no competitor holds. That is the part that matters for AI discovery. Your competitor can read your product page, your category, your reviews and your competitors’ pages. They cannot read your returns log. A page written from it carries something theirs cannot.

Ten codes at SKU resolution, aggregated across a season, will tell you more about which pages are failing than any keyword tool, because a keyword tool describes a market and a returns log describes your customers meeting your product.

Turning it into pages

The workflow is unglamorous.

Export returns for the last two quarters with the reason code and the SKU. Group by product, then by reason. Ignore anything with one or two returns; you want products where the same reason repeats, because repetition is the signal that the page is at fault rather than the individual buyer.

For each repeat, write down the fact that would have prevented it. A returned garment marked “size too small” three times needs the actual measurements and a comparison to a familiar reference, not a size chart link. A product returned as “not as described” needs the specific attribute the description over-promised, corrected.

Then put that fact on the page as visible text. Not in a size chart popup, not in a downloadable spec sheet, not in a tab that loads on click. Assistants read what is served in the page. A fact that appears only after a JavaScript interaction is a fact they will not see.

If the same missing fact turns up across a whole range, it belongs at collection level as well as on each product.

What this does not prove

Two limits worth stating.

Nobody has shown that adding these facts makes an assistant recommend you. There is no study on real product data and commercial assistants, and anyone claiming a percentage lift from it is inventing one. The argument here is narrower: a page that answers a question a real buyer got wrong is more useful to a machine than one that does not, and the fact came from a record only you hold.

And returns data has a survivorship problem. It only tells you about buyers who bought. The shopper who read the page, could not find the measurement and left is invisible in it. Your on-site search log is closer to that person, which is why Search Console cannot see your AI discovery problem argues for reading zero-result queries alongside this.

The test to apply first

Before you write any of these pages, run one check on the brief.

Could a competitor with access only to public sources have written this page? For a page built from your returns log, the answer is no, and that is the whole point. For most of the content in most ecommerce calendars, the answer is yes.

A page that fails that test may still be worth publishing. It will not be worth citing.

For why understandable beats impressive when a machine is choosing, see AI recommends what it can understand. For the difference between depth and coherence, see SEO rewards depth, AI rewards coherence.

Frequently asked questions

Where do I find return reasons on Shopify?

Return reasons live on return line items in the Admin API, using the `ReturnReason` enum. There is also a `returnReasonNote` field capped at 255 characters and a customer note capped at 300. Merchants on a third-party returns platform may have richer free text — Loop Returns, for example, carries a free-text reason in its warehouse reporting endpoint.

Our return rate is low. Is this still worth doing?

Yes, and it is quicker. A low return rate means a short list, and you are reading it rather than counting it. Ten repeated returns against one product is a clear enough instruction to rewrite that page.

Should we publish the return reasons themselves?

No. Publish the fact that would have prevented the return — the measurement, the material, the compatibility note. The return is the source, not the content.

Does this apply to WooCommerce?

The principle does. The plumbing differs: WooCommerce has no native returns object, so the data sits in whichever returns or RMA plugin the store runs. Check what your plugin exposes before planning work around it.

Is this the same as reading reviews?

Related but not the same. Reviews are usually public, so a competitor can read them too. Returns data is not public, which is exactly what makes a page written from it hard to reproduce.

Next step

Ready to apply this to your catalogue?

Move from theory to implementation with parity-first GEO workflows.

Turn your catalogue into AI-readable product pages and structured data.

Already have an account? Sign in

Structured outputs enabled
First-party pages published
Discovery coverage expanding