How we get to the number, published in full
This page is the arithmetic behind every figure on this site: how the sample is built, how many answers it contains, how the interval around the number is computed, what size of change we are willing to call real, and what we do not measure at all.
Some of what follows makes our own numbers look less impressive than they would without it. That is the point. A measurement you cannot check is not a measurement, it is a claim.
What is being measured
The headline figure is one thing and nothing else: the share of model answers that mention your brand, measured over a fixed set of questions.
The questions are generated from your own site — your brand, your products, your region, your competitors. Once generated, they are held fixed. You cannot edit them, and neither do we between measurements: a question pool that drifts makes two weeks incomparable, and every delta computed across it would be meaningless.
The rest of the report comes out of the same answers, not out of a second source: which brands are named instead of you, and where a model asserts something your own site contradicts.
Why we ask the same question more than once
A model asked the same question twice can answer differently. One answer is therefore not a measurement — it is a single draw from a distribution, and a distribution is what we are trying to describe.
So every measurement is a sample, and the repeats are what make it one:
- Free check: 3 models × 5 prompts × 3 repeats = 45 answers.
- Paid audit: 4 models × 15 prompts × 5 repeats = 300 answers.
Every answer is one observation. Without the repeats there would be nothing to compute an interval from, and the number would be exactly the kind of single figure this page exists to argue against.
The interval, and the confidence level it is computed at
The share of answers mentioning your brand is a binomial proportion over n independent answers. The half-width of the interval around it is approximately 1.96 × √(p(1−p)/n).
Every interval and every threshold on this site is computed at 95% confidence. When we print "95% interval 18–44%", we mean that an interval built this way contains the true share in 95 measurements out of 100.
We evaluate the width at p ≈ 0.3, the least precise point in the range we actually work in. The interval we publish is therefore the wide case, not the flattering one.
| Answers in the sample | Half-width of the 95% interval | Smallest change between two measurements that means anything |
|---|---|---|
| 45 — free check | ±13.4 pp | 18.9 pp |
| 100 | ±9.0 pp | 12.7 pp |
| 300 — paid audit | ±5.2 pp | 7.3 pp |
| 480 | ±4.1 pp | 5.8 pp |
The two rows you are not sold are in the table as well. They are how we chose the two you are.
When we will say a change is real, and when we will not
Two measurements taken with the same method still differ by chance. So a difference only means something once it is larger than the threshold for its sample size:
- Free check, 45 answers: a change smaller than 18.9 pp means nothing.
- Paid audit, 300 answers: a change smaller than 7.3 pp means nothing.
Below the threshold we mark the change indistinguishable from noise — including the weeks when it moved in the direction you were hoping for, and including the weeks when the honest answer is that nothing measurable happened.
This is also the whole difference between the free check and the paid audit. You are not buying more features. You are buying sample size, and sample size is what turns a number into a change you can act on.
Why 45 and 300
300 is the first sample size at which the significance threshold drops below 8 pp while the audit stays cheap enough to re-run every week. Doubling it to 480 improves the threshold by 1.5 pp and costs 60% more compute for it — that trade does not pay for itself, and you would be the one paying for it.
45 is chosen for the opposite reason. It is enough to show you the picture, and it is not enough to track a change. We would rather tell you that in the free result than let you read movement into noise.
What this measurement is not
We query models over their APIs. That is not the AI Overviews block in Google search, and it is not what you see in the ChatGPT web app or in any other assistant's web interface. Those surfaces have their own retrieval, their own personalisation and their own ranking on top of the model. Our numbers do not describe them, and we do not claim they do.
Our questions are generated from your site, not taken from real user traffic. We do not know what people actually type into assistants, and nothing here should be read as query volume.
We do not measure whether an edit to your site caused a change. We measure before, we measure after, and we tell you whether the difference is bigger than noise. A causal link between an edit and a model's output is not externally provable — by us or by anyone selling you the opposite.
What happens when a measurement is imperfect
If some answers do not come back, we do not quietly top the sample up. The measurement runs on the answers we actually received, the interval is computed on that number, and the report carries the line "measured over N answers instead of 45".
The free check runs once per domain every 7 days. Ask sooner and you get the report we already ran, not a new one: a fresh 45-answer measurement on the same day would differ from the first by chance alone, and we are not going to hand you that as news.
Two free checks a week apart are comparable only if they differ by more than 18.9 pp. The email offering you the repeat says the same thing.
Every figure on this site is produced exactly the way this page describes. Read the interval before you read the number — that is why both are printed.
Run the free check