Every ecommerce team asking about GEO eventually asks the same question: how do I know if it’s working?
It’s the right question, and most of the industry answers it badly. So let’s do this properly: what you can measure, what you can’t, and how to tell an honest measurement system from a confident-sounding one.
The problem is actually quite simple
When a shopper asks an AI assistant for a recommendation, the assistant names a few products and moves on. No impression data reaches you. No search console. No rank tracker. The most important new surface in product discovery ships with no analytics.
So measurement has to be reconstructed from the outside: ask the engines the questions your customers ask, record the answers, and extract the structure from them. That’s straightforward to do badly and hard to do well — and the difference is worth understanding before you trust any number.
What’s actually worth tracking
1. Presence and rank. For a defined set of high-intent buyer questions — “best sustainable leather boots under £150?”, not vague keywords — are you named? In what position? Naming order matters: assistants front-load their confidence.
2. Share of voice. Across your question set, how often are you named versus the competitors who keep appearing? The competitor list the engines produce is itself intelligence — it rarely matches the competitor list in your head.
3. Citations — who the engine trusted. When an answer cites sources, which won? Your product page, a retailer listing, a review site, a video? This tells you where the answer is actually decided, and it’s frequently not where you’re spending effort.
4. The pathway — the diagnostic almost everyone skips. There are two very different ways an engine produces an answer. It can retrieve — fetch pages live and synthesise from them. Or it can answer parametrically — from what it absorbed in training. The distinction decides your entire remediation strategy. Retrieved-but-not-chosen: your page was read and lost — fix the page. Parametric: the engine’s memory of your category doesn’t include you — no page edit fixes that; you need authority on the sources engines learn from. Same gap, opposite fixes. Measurement that doesn’t tell you the pathway tells you that you lost, never why.
5. Movement against a control. AI answers drift on their own — models update, competitors act. If you change nothing and visibility moves anyway, that’s the noise floor. Movement only means something measured against a baseline of untouched questions and benchmark brands.
What an honest instrument refuses to do
The uncomfortable part: this is probabilistic measurement of non-deterministic systems. The same question can produce different answers an hour apart. Any system claiming certainty here is overclaiming — and you’re going to spend budget on these numbers, so overclaiming isn’t a cosmetic sin.
Four behaviours to demand from anything you use — including ours:
- Zero vs unknown. “We measured, you weren’t named” and “we couldn’t measure this” are different facts. A system that renders both as zero is fabricating data.
- Estimated vs verified, labelled. Some results come from APIs (fast, broad, approximate), some from checking the live product surface (slower, truer). You should always know which you’re looking at.
- Correlation, not causation. “You changed X and visibility rose” is a correlation claim. Genuine causal proof needs holdouts and time. A system that says “we improved your visibility by 40%” without a control is marketing, not measurement.
- Abstention. Sometimes the honest answer is “we don’t know yet”. A system that never says this is guessing somewhere.
We built Geoffy Monitor around exactly these rules — permanent holdout, benchmark drift basket, zero-vs-unknown discipline, labelled fidelity. Not because it makes the numbers more impressive; because it makes them safe to act on. Monitor is in early access with a small group of customers now. Request early access →
Start smaller than you think
You don’t need a thousand tracked prompts. You need thirty good questions — the ones with a buyer behind them — measured consistently, with the pathway split, against a baseline. That beats a wall of vanity dashboards every time.
And if you just want to know where you stand today: the free GEO Score runs this measurement once, on your real products. It’s the same instrument, as a snapshot. Get your free GEO Score →