Method
How we measure something that changes every time you ask it.
AI answers are non-deterministic, personalised and geography-dependent. That makes honest measurement harder. It does not make it impossible — it makes the method the product.
End to end
From a URL to a decision.
This is the same pipeline whether it runs once as an audit or monthly as part of the growth engagement. The difference is cadence, not method.
- 01
Understand the company
crawl + interviewA full crawl, plus a session with someone who knows the product. We need to know what you actually sell, to whom, in which markets, and what a buyer would type — or say — when looking for it. - 02
Build the query set
generated, then agreedQueries derived from industry, product, customers, geography, use cases, competitors and intent, spanning the archetypes that matter in a considered purchase: best X, best X for Y, alternatives, versus, competitors, who provides X, enterprise X, affordable X, how it works, and what it costs. You see the set and sign off before anything runs. - 03
Establish the competitor set
four sourcesWho you say your competitors are, who the market says, who ranks against you, and who the answer engines actually name. Share of AI visibility is computed against this set, so the set is agreed and recorded rather than assumed. - 04
Execute with repetition
the part most people skipEvery query, on every surface, several times. Raw responses stored verbatim alongside the model identifier, timestamp, geography and settings. A single result is an anecdote; a rate over repetitions is a measurement. - 05
Extract mentions and citations
with false-positive controlWas the company named, in what position, framed how? Was its own domain cited, or was a third party cited that happens to mention it? Brand-name variants are matched against an allow list, and low-confidence matches go to a human. - 06
Trace the evidence chain
the diagnostic that mattersFor every query where a competitor wins, collect the sources the answer used and read what they say. This is how a vague 'they have more authority' becomes a specific list of pages you have no equivalent of. - 07
Check technical accessibility separately
different problemCan these systems even reach you? Per-bot robots directives, edge and WAF behaviour, rendering. Being invisible because you are blocked and being invisible because you are unconvincing require opposite responses. - 08
Prioritise and recommend
human ownedFindings become recommendations carrying problem, evidence, hypothesis, expected impact, effort, priority, affected queries, affected competitors, owner, and whether the work is automatable. - 09
Ship, then re-measure
same methodWork is timestamped on the timeline. The next period runs the identical query set so any movement can be attributed to something specific rather than narrated after the fact.
Metrics
Six definitions, no proprietary score.
Each metric is defined precisely enough that two people computing it independently get the same number. Where a definition involves a judgement call — what counts as a relevant query, who is in the competitor set — the judgement is written down, not buried.
Mention rate
mentions ÷ relevant queries
The headline number. Reported per surface, never averaged across them.
Citation rate
queries citing your domain ÷ relevant queries
Different from mention rate, and often much lower. Being named is not being sourced.
Average mention position
mean ordinal position when mentioned
With explicit rules for prose answers that have no meaningful order.
Query coverage
queries with any appearance ÷ query set
Breadth rather than depth. Catches the company that dominates three queries and is absent from thirty.
Competitor visibility
the same metrics, per competitor
Same query set, same executions, same method. Comparable by construction.
Share of AI visibility
your appearances ÷ appearances across the set
Meaningless without the competitor-set definition, so the definition travels with the number.
A composite “AI score”
A single number out of 100 is easy to put on a slide and impossible to act on. Worse, it lets a vendor move the weighting until the chart looks like progress. If we ever publish a composite, the full formula and weights get published with it, and it will never replace the underlying rates.
Provenance
Every number says where it came from.
Six labels, applied to every quantitative claim in every deliverable. This is enforced in the product's type system, not left to the discipline of whoever is writing the report at midnight.
- Observedstatable as fact
- We ran the check ourselves and recorded the raw result. The underlying response, header or HTML is stored and can be replayed.
- Verifiedstatable as fact
- Observed, and then independently confirmed a second time — by a repeat run, a second source, or a human review. The strongest label we issue.
- Inferred
- A conclusion drawn from observed evidence rather than directly measured. The evidence chain is recorded alongside it.
- Estimated
- A modelled or sampled figure with known error bars. Never presented without stating the basis of the estimate.
- Mock
- Hand-authored placeholder data used to build and test the product. Never valid in a client deliverable.
- Demo
- Illustrative data shown publicly to explain what the product does. Always labelled as such on screen, and about a company that does not exist.
Connecting it to money
How far down the funnel we can honestly go.
We measure as far as your stack allows and then stop, rather than modelling past the edge of the data and presenting the model as a result.
- DiscoverabilityDirectly measured
- UnderstandingDirectly measured
- AuthorityDirectly measured
- AI visibilityDirectly measured
- Website visitMeasured where referrers survive
- LeadYour analytics and CRM
- CustomerYour CRM
- RevenueAttribution permitting
The attribution problem, stated plainly
A buyer who reads an AI answer naming you, then searches your brand and arrives directly, appears in your analytics as direct traffic. That influence is real and largely invisible. We will not invent a multiplier to account for it. What we do instead is track brand search volume and direct traffic alongside visibility, and let you draw your own conclusions from the correlation.
Limits
What this method cannot do.
- It cannot observe personalised answers. We measure clean sessions, which is a reasonable proxy for a new buyer and a poor one for an existing user.
- It cannot see inside any provider’s selection logic. Everything we say about why something happened is an inference from observed evidence and is labelled as one.
- It cannot fully separate your work from the model changing underneath you. More periods of data narrow this; nothing eliminates it.
- It cannot give you a guarantee. If that is the requirement, we are the wrong choice, and so is anyone who says yes.
See where you stand in AI search.
An AI Search Audit tells you how often AI systems name your company, who they name instead, and what is causing the gap. Every figure comes with the method behind it.