Blog

Measuring AI discovery: what to track, and what to refuse to claim

Most AI visibility numbers are theatre. Here's what's actually measurable in AI product discovery, what isn't, and how an honest measurement system behaves.

By Anthony Gale — Co-Founder, Geoffy

AI DiscoveryMeasurementGEO
Measuring AI discovery: what to track, and what to refuse to claim cover image

Playing the video loads it from YouTube, which sets its own cookies. The full transcript is below, so you can read it without doing that.

Read the full transcript

Share of voice in AI answers is the share of relevant buying questions where the assistant names your brand — measured across the engines that matter, run repeatedly, tracked over time. It's the single number that tells you whether AI is recommending you or your competitors. And here's why "measured properly" matters: ask the same assistant the same question twice and you can get two different answers. A one-off check isn't a measurement — it's a coin flip you happened to watch once.

First, why this number and not "do I appear in ChatGPT?" Appearing once is an anecdote. Share of voice is a rate: out of all the questions a buyer might ask in your category, in what share are you named — and how does that compare to each competitor? That turns a vague worry into something you can manage, set targets against, and watch move when you make changes. And it's a number worth managing, because every point of share of voice is a slice of the buyers who never see a results page at all.

So how do you build it? Four ingredients.

One: the right questions. Not your brand name — the questions a first-time buyer asks before they know you. "Best X for Y." "Affordable X that does Z." You need a real set of them, covering how shoppers actually phrase the need. Get the questions wrong and the whole measurement is wrong.

Two: every engine that matters. ChatGPT, Perplexity, Google's AI Overviews and AI Mode, and the others your buyers use. They don't agree — you can be the top recommendation in one and absent from another. A single-engine number hides half the picture. Each engine is a separate shelf, and share of voice has to be measured shelf by shelf.

Three: repetition. This is the one people skip, and it's the one that matters most. You don't have a ranking in ChatGPT. You have a probability — and probabilities can only be measured by counting. That's the coin flip from the top: run each question many times and count the mentions as a rate, or you're not measuring anything.

Four: tracking over time. A share-of-voice number on its own is a snapshot. The value is in the trend — did it move after you fixed your product pages, published that answer, earned that review? Without a time series you can't tell whether your work is paying off or you're standing still while competitors climb.

Now put those four together and you'll see the problem: the right questions, times every engine, times many repeat runs, tracked every week. That's hundreds of queries on a schedule. It is not a job you do by hand on a Tuesday — by the time you'd finished one round, it'd be out of date.

That's precisely what Geoffy automates. You give it your category and competitors; it runs the real buying questions across the major AI assistants on a schedule, records every time you're named, who's named instead, and where you rank — and turns all of it into one visibility score and a trend line. But counting mentions is table stakes — plenty of tools do that. The score exists to prove your fixes worked, which is why Geoffy pairs it with the modules that actually engineer your product pages and build your off-site corroboration — so the number you're watching is the number you're moving.

If you want the manual version first, the earlier video on checking whether ChatGPT recommends your brand walks through it. If you want the number this week — and the fixes that move it — that's geoffy dot ai. See you in the next one.

Every ecommerce team asking about GEO eventually asks the same question: how do I know if it’s working?

It’s the right question, and most of the industry answers it badly. So let’s do this properly: what you can measure, what you can’t, and how to tell an honest measurement system from a confident-sounding one.

The problem is actually quite simple

When a shopper asks an AI assistant for a recommendation, the assistant names a few products and moves on. No impression data reaches you. No search console. No rank tracker. The most important new surface in product discovery ships with no analytics.

So measurement has to be reconstructed from the outside: ask the engines the questions your customers ask, record the answers, and extract the structure from them. That’s straightforward to do badly and hard to do well — and the difference is worth understanding before you trust any number.

What’s actually worth tracking

1. Presence and rank. For a defined set of high-intent buyer questions — “best sustainable leather boots under £150?”, not vague keywords — are you named? In what position? Naming order matters: assistants front-load their confidence.

2. Share of voice. Across your question set, how often are you named versus the competitors who keep appearing? The competitor list the engines produce is itself intelligence — it rarely matches the competitor list in your head.

3. Citations — who the engine trusted. When an answer cites sources, which won? Your product page, a retailer listing, a review site, a video? This tells you where the answer is actually decided, and it’s frequently not where you’re spending effort.

4. The pathway — the diagnostic almost everyone skips. There are two very different ways an engine produces an answer. It can retrieve — fetch pages live and synthesise from them. Or it can answer parametrically — from what it absorbed in training. The distinction decides your entire remediation strategy. Retrieved-but-not-chosen: your page was read and lost — fix the page. Parametric: the engine’s memory of your category doesn’t include you — no page edit fixes that; you need authority on the sources engines learn from. Same gap, opposite fixes. Measurement that doesn’t tell you the pathway tells you that you lost, never why.

5. Movement against a control. AI answers drift on their own — models update, competitors act. If you change nothing and visibility moves anyway, that’s the noise floor. Movement only means something measured against a baseline of untouched questions and benchmark brands.

What an honest instrument refuses to do

The uncomfortable part: this is probabilistic measurement of non-deterministic systems. The same question can produce different answers an hour apart. Any system claiming certainty here is overclaiming — and you’re going to spend budget on these numbers, so overclaiming isn’t a cosmetic sin.

Four behaviours to demand from anything you use — including ours:

  • Zero vs unknown. “We measured, you weren’t named” and “we couldn’t measure this” are different facts. A system that renders both as zero is fabricating data.
  • Estimated vs verified, labelled. Some results come from APIs (fast, broad, approximate), some from checking the live product surface (slower, truer). You should always know which you’re looking at.
  • Correlation, not causation. “You changed X and visibility rose” is a correlation claim. Genuine causal proof needs holdouts and time. A system that says “we improved your visibility by 40%” without a control is marketing, not measurement.
  • Abstention. Sometimes the honest answer is “we don’t know yet”. A system that never says this is guessing somewhere.

We built Geoffy Monitor around exactly these rules — permanent holdout, benchmark drift basket, zero-vs-unknown discipline, labelled fidelity. Not because it makes the numbers more impressive; because it makes them safe to act on. Monitor is in early access with a small group of customers now. Request early access →

Start smaller than you think

You don’t need a thousand tracked prompts. You need thirty good questions — the ones with a buyer behind them — measured consistently, with the pathway split, against a baseline. That beats a wall of vanity dashboards every time.

And if you just want to know where you stand today: the free GEO Score runs this measurement once, on your real products. It’s the same instrument, as a snapshot. Get your free GEO Score →

Next step

Ready to apply this to your catalogue?

Move from theory to implementation with parity-first GEO workflows.

Turn your catalogue into AI-readable product pages and structured data.

Already have an account? Sign in

Structured outputs enabled
First-party pages published
Discovery coverage expanding