Skip to main content

Measurement

How we measure

We measure by hand, not through APIs.

The setup

The setup is always the same: ten buyer questions that do not contain your brand name, across four engines, ChatGPT, Perplexity, Google AI Overviews and Claude, with three runs per question and engine. That is 120 documented data points per domain and measurement round. Every run happens in its own fresh session, as a reference measurement without personalisation, fixed region Germany.

The scoring

Every run is scored individually from 0 to 3: not mentioned, mentioned, cited as a source, actively recommended. Scoring happens live during the measurement, and we file dated screenshots for the core findings. From 16 September on we document every single run with a screenshot.

Google does not return an AI Overview for every question. From 16 September on we record separately whether an answer appeared at all, and we report two numbers for Google: how often an Overview appeared and how often the brand appeared in it. Otherwise a question without an Overview would look exactly like a question where the brand was missing.

Two measurement points

We measure twice with an identical question set, once before the work and once after. Declines are reported, they go in the report.

Why not automated

We measure the surface your buyers search in, not the interface behind it. For ChatGPT and Perplexity the API answers differently from the interface: different system prompt, different retrieval mechanics, different number. For Google AI Overviews there is no official interface, and the available scrapers read the search results page. We deliberately switch personalisation off, otherwise two measurement rounds would not be comparable.

Effort and verifiability

The effort is around five hours per domain and measurement round: two and a half hours of measurement, one hour of evaluation, one and a half to two hours for the report. That is expensive and slow. In return you get the complete question set and can ask every question yourself. Individual runs fluctuate, which is why we measure every question three times. If your sample differs from our result, we will show you the evidence.