There’s a category of Shopify store right now doing all the visible AI-visibility work — schema, intent-layered content, comparison pages, clean copy — and still getting zero citation hits. Not because the strategy is wrong. Because the plumbing is.
The one-click own-goal
If your store sits behind Cloudflare and you’ve enabled “Managed robots.txt” — a single toggle in the AI Crawl Control panel — Cloudflare blocks a default list of AI crawlers on your behalf. That sounds reasonable. The problem is that the list lumps together two very different categories of bot.
Training crawlers scrape pages to feed back into model training: GPTBot, CCBot, Bytespider, Amazonbot, meta-externalagent. Blocking these is the standard “I don’t want my content training someone else’s foundation model” posture, and it’s defensible.
But the same managed list also disallows ClaudeBot — a crawler that does fetch pages to support Claude’s answers — alongside the AI-training control tokens Google-Extended and Applebot-Extended. So a toggle most merchants read as “don’t train on my content” also removes one of the live fetchers that can cite you.
Other citation fetchers, including PerplexityBot and OAI-SearchBot, are not on Cloudflare’s managed list at all — you have to check your own robots.txt for those.
The practical consequence is the same either way: you can spend months on structured data and answer-shaped content and still be missing from AI answers, because the fetcher gets a 403 when it tries to read the page. Not because anyone made a deliberate choice, but because the toggle was easy and the distinction wasn’t.
The fix: split the posture
Decide the training question and the citation question separately, rather than letting one toggle answer both.
- Disallow, if you don’t want your content training foundation models: GPTBot, CCBot, Bytespider, Amazonbot, meta-externalagent
- Reconsider blocking: ClaudeBot, which fetches pages to support answers
- Decide deliberately: Google-Extended and Applebot-Extended are training opt-out controls, not citation fetchers. Google states plainly that Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal” — so allowing it is a decision about model training, not about visibility
- Check separately: PerplexityBot and OAI-SearchBot aren’t on the managed list, so your own robots.txt governs them
You keep the “don’t train on my content” stance — which is what most brands actually mean when they say they don’t want AI using their content — without quietly closing the door on the answer surfaces you want traffic from.
One currency note: from 15 September 2026 Cloudflare splits crawlers into Search, Agent and Training categories, with search crawlers remaining allowed by default. Re-check your posture after that lands.
Two related findings worth knowing
The Merchant Center → Gemini route is muddier than it looks. The intuitive assumption — “our clean product feed flows to Google Merchant Center, surely that gets us cited in Gemini’s chat answers” — doesn’t map onto reality cleanly. The feed powers the Shopping graph, which surfaces through search results, Google Shopping and the shopping carousels inside AI Overviews. It’s not the same retrieval flow as Gemini’s conversational citations. Feed quality lifts your odds in the shopping module; it doesn’t necessarily get you mentioned in the prose of a chat reply.
Comparison-shape content beats list-shape content for citation. “X vs Y: when to pick each” outperforms “10 best alternatives to X” by a wide margin. The mechanism is question-shape matching: a comparison page gives the model a paragraph already structured as a defensible answer to “should I buy X or Y?”, so it can cite cleanly. A list post forces the model to summarise the whole page or pick one item and lose context — and it tends to do neither well. Granularity compounds this: “magnesium glycinate vs citrate for sleep” beats “10 best magnesium supplements”, because the narrower comparison maps to a more specific question.
Why this matters more than it sounds
AI search is currently rewarding hygiene work that almost nobody is doing. The flip side is the opportunity: if your competitors have left the Cloudflare default on, you can pass them on visibility for the price of editing one file. That’s not a normal SEO situation. Normal SEO doesn’t have a setting that hides you from half the answer engines — one that nobody remembers enabling.
If you run a Shopify store, open your robots.txt now and check who’s allowed. If you’re behind Cloudflare, check the AI Crawl Control panel too. The five citation bots above are the ones to allow if you want any chance of appearing in AI answers. For the rest of the discovery stack — llms.txt, structured data, intent-tier content — see how Geoffy prepares Shopify stores for AI discovery.
Not sure what’s blocked on your store? Get your free GEO Score — crawler access is one of the first things it checks.