Skip to content
All insights
Measurement30 July 2026 · 8 min read

Measuring AI visibility without fooling yourself

Non-determinism, tiny denominators and moving query sets make it trivially easy to produce numbers that feel rigorous and mean nothing. Six rules we hold ourselves to.

Measuring visibility in a system that returns a different answer each time you ask is genuinely hard. It is also easy to do badly in a way that looks completely convincing on a slide, which is the more dangerous outcome.

1. Never report a rate without its denominator

Twenty-five percent is four queries with one hit. It is also four hundred queries with a hundred. Printed as a percentage they are indistinguishable, and one of them is noise. Every rate we publish carries its numerator and denominator in the same visual unit, and our report generator treats a missing denominator as a build failure rather than a style issue.

2. Repeat every query, and count executions

A single run tells you what one sample of a stochastic process produced. Repetition is not optional rigour, it is the minimum condition for the number to mean anything. It also costs money, which is the real reason it gets skipped.

How many repetitions is a budget decision. What matters is that the count is fixed, disclosed, and held constant between periods, because changing it changes the metric.

3. Never average across surfaces

ChatGPT, Gemini and Perplexity behave differently, draw on different sources and change on different schedules. A combined figure hides the only interesting part: where you are strong, where you are absent, and whether the gap is about access or about evidence.

4. Establish a noise band before you interpret movement

On a 40-query set, one query flipping is 2.5 percentage points. Movement of that size between periods is not a result. Decide in advance, from repeated baseline runs, how much variation the system produces when nothing has changed — and then report anything inside that band as noise, including when it is favourable.

The uncomfortable version of this rule

It applies in both directions. The temptation to explain a small favourable movement is much stronger than the temptation to explain a small unfavourable one, and giving in to it is how measurement becomes marketing.

5. Freeze the query set, and flag it when you cannot

Period-over-period comparison assumes the same set of questions. Adding queries because a new use case emerged is legitimate and necessary; doing it silently breaks every trend line that crosses the change.

We store a flag on each snapshot recording whether the query set changed, and the report says so explicitly rather than absorbing the change into a chart.

6. Separate what you saw from what you think it means

The response is an observation. Why the model produced it is an inference. Whether a fix will change it is a prediction. These get different labels in every artefact we produce, because the moment they blend, a plausible story becomes indistinguishable from a recorded fact.

On composite scores

A single number out of a hundred is easy to put on a board slide and nearly impossible to act on. It also gives whoever owns the weighting a quiet lever over whether the chart shows progress.

We do not publish one. If we ever do, the formula and the weights get published with it, the inputs stay visible underneath, and it will never be the number we ask a client to make a decision on.

See where you stand in AI search.

An AI Search Audit tells you how often AI systems name your company, who they name instead, and what is causing the gap. Every figure comes with the method behind it.