How we measure.
Most tools in this category won't tell you how often they ask, or how far a result drifts on its own. A couple of the big enterprise ones do, and do it well. Almost nobody at this price does.
We think it's the most important thing about a number, so here's ours in full.
What we ask
Five questions. The ones your buyers would actually type at the moment they're choosing someone, not brand searches.
You approve them before anything runs. If you read the five and think nobody would ever ask that, we go again. The whole exercise is worthless if you don't believe the questions.
The questions are written for a specific market, because the answers genuinely differ by country. You set the market when you sign up, and you can go narrower than a country if your buyers are local.
Where we ask
Five surfaces, automatically, every month, in the same way every time.
ChatGPT, Gemini, Grok, Perplexity and Google AI Overviews.
And two more by hand, once a quarter: Claude and Microsoft Copilot.
Those two are separated for a reason. They can't be asked automatically without swapping in a different system and calling it their answer. An API answer is not the answer a person gets when they open the app, and we won't print one and label it the other. So we ask those two ourselves, and we tell you that's what we did.
Where we ask from
If your buyers are in one town rather than one country, your five questions name the town, and every surface is asked about it in the words you approved.
Three of the five also take a location directly, and are searched from your area: ChatGPT, Perplexity and Google's AI answers. Grok and Gemini accept no location parameter of any kind, from anybody, so for those two your area reaches the answer through the question and nothing else. Every report says which surfaces got which.
We hold the network we ask from constant, in the same place, every month. It is the only geography Grok and Gemini have, and a network origin that wandered would move your number for a reason that has nothing to do with your market.
How many times we ask
55 captures per report.
Gemini, Perplexity and Google AI Overviews get asked three times per question and the result is averaged. ChatGPT and Grok get asked once.
That difference isn't arbitrary. It follows what each one costs to ask, not how much it matters, and we would rather say that than imply the five are sampled evenly. The three we repeat are the three we can afford to repeat.
What we do with the answers
Two separate steps, and keeping them separate is the point.
First we capture. The full answer, word for word, exactly as it came back, stored with the date, the market, and which system gave it.
Then we read. Was your brand named at all. Was it recommended, or just listed. Which competitors were named. Which sites the answer drew from.
We keep the raw answers because the reading is an opinion and the answer isn't. You get both, so you can disagree with us.
The number on the front page
Direct Endorsement: how many of the surfaces that answered recommended you, out of the surfaces that answered.
The denominator is the surfaces your report actually covered that month, not a fixed five. A surface that produced no answer leaves the count rather than scoring as a no, and on a Plus quarter month the two we read by hand join it once they have been read. A count is only worth anything if you know what it is out of, so the report prints both numbers and never a percentage.
Not how many mentioned you. Named and recommended are different things, and the gap between them is usually the story.
Recommendation Share is measured on the four questions that name no business, over the five monthly surfaces.
That is the line on the chart, and it is a different number from the one above it. The fifth question asks about you by name, so on it you are the subject and a competitor is at best an alternative — counting it would lift your line for a reason that is the question rather than your market. Claude and Microsoft Copilot stay off the line too: they are read four times a year, and folding them in would move it every third month because of our schedule.
Where a surface is asked three times, it counts as recommending you if most of those readings did, not if any one of them did. One reading in three saying yes is not a surface that recommends you, and counting it that way would round every borderline case in our favour.
It's a status, not a trend. There's no arrow on it, ever, and we measured why. Read on.
Underneath it we show presence, which is how often you were named at all across all the questions, and we state the gap between the two in plain words.
The bit nobody else publishes
Ask an AI the same question twice and you can get different companies. So before we reported that something changed for you, we needed to know how much these systems move on their own.
So we measured it. Same question, same surface, same market, ten times back to back, with nothing changed in between. Real questions from a real subscriber's scope.
| asked | brand named | competitor list unchanged | cited sites unchanged | |
|---|---|---|---|---|
| Gemini | 10 of 10 | 4 of 10 | 87% | 19% |
| Perplexity | 10 of 10 | 0 of 10 | 56% | 29% |
| ChatGPT | 10 of 10 | 0 of 10 | 63% | 32% |
Measured by us, 23 August 2026. Naming, on questions that don't mention you.
A brand sitting well outside the answer reports stably. Twice at zero out of ten. But a brand sitting near the edge of being mentioned is a coin flip, named four times out of ten with nothing whatsoever happening in the market.
The edge is exactly where a business making progress lives. Which means the instrument is least reliable precisely where it matters most to you.
And the competitor list moves even when your own mention doesn't. Between a third and nearly half of the named companies change between identical runs.
And the same test on the number at the top
The table above counts whether a brand was named. The number on the front page is whether a surface recommends you when somebody asks about you directly. Those are different questions, so we ran the same test on the second one. Ten readings, each surface, asked about a real brand by name, nothing altered in between.
| answered | named you | recommended you | |
|---|---|---|---|
| ChatGPT | 9 of 10 | 9 of 9 | 0 of 9 |
| Gemini | 10 of 10 | 10 of 10 | 8 of 10 |
| Grok | 10 of 10 | 10 of 10 | 0 of 10 |
| Perplexity | 10 of 10 | 10 of 10 | 0 of 10 |
| Google AI Overviews | 10 of 10 | 10 of 10 | 0 of 10 |
Measured by us, 25 August 2026. Recommending, on the question that names you.
One surface in five changes its verdict on its own. That is a fifth of the number on your front page moving with nothing happening in your market, which is why it never carries an arrow. Four of the five said the same thing ten times out of ten.
It is the same shape as the naming result. A brand sitting firmly outside the recommendation gets a stable no. The one surface near the boundary is the coin flip. Which means a business whose score is 1 out of 5 may be resting that whole 1 on the least reliable reading in the set, and nothing on the face of the number says so. So we say it: where a surface didn't give the same verdict every time, your report prints how the readings split.
How much of that drift is ours
What you read is the surface's answer plus our reading of it, and our reading is a language model too. A model set to be as repeatable as possible is not the same as one that is.
So we checked. We took the first answer from each surface and read it three times over, unchanged. Fifteen readings of identical text, and not one disagreement.
The drift in the table above is the surfaces. If it had been ours, that would be a defect for us to fix, not a floor for you to work around, and we would have said that instead.
What we still haven't measured
Both tests were run on a national scope. If your buyers are in one town, the answer sets are thinner and may drift further, or less. Nobody has measured that, including us, so nothing on this page claims anything about it.
It is one command and a few dollars, and it will be here when it is done, whichever way it comes out.
What we do about it
We don't report a change we can't distinguish from noise.
We know how far a result drifts on its own, so anything smaller than that reports as steady. Not as an improvement, not as a decline. Steady. Some months your report will say nothing moved further than we can measure, and that will be the honest answer.
If we ever change how a number is calculated, we say so and we stop comparing across the change rather than quietly carrying on. The month a definition changes is a break, and breaks get declared.
Why we don't check every day
Four findings, none of them ours, all from research the tool vendors published themselves.
- Semrush, across 1,094 categories and 600,000 citations: the brand that owned a topic held the number one spot in 90.4% of month-over-month comparisons. The leaders are stable.
- Zatuchin (arXiv:2607.13304), across 12,933 responses: brand identity accounts for 1.5% of the variance in a single answer. The language you ask in accounts for 26.5%. Most daily movement is a phrasing artefact.
- Conductor, across 13,770 domains and 3.3 billion sessions: AI referral traffic is 1.08% of website traffic. BrightEdge independently puts it under 1%.
- Evertune: a single prompt asked once carries plus or minus 44 points of error. Which cuts both ways. It is why a daily number is noise, and it is why nobody should sell a precise one.
Those four are other people's measurements and we have named them so you can go and read them. Everything else on this page is ours, and you can check it against your own report.
A source that keeps appearing across 55 captures is real. A competitor named again and again is real. What the answer actually said is not a statistic at all, it is evidence. None of those need the sample size a percentage needs, which is why the report leads with what was learned rather than with what moved.
What we don't claim
We don't claim a person reads every answer every month. Five surfaces are captured and read by the system. Claude and Copilot are read by a person, every three months from the date a subscription starts. That's the whole of it.
We don't claim to measure surfaces we don't measure. The report names which ones it covers, every time.
We don't claim precision we haven't got. The table above is on this page permanently, including the parts of it that are inconvenient for us, and so is the section saying what we have not measured yet.
Why we bother publishing this
Because the alternative is asking you to trust a number with no working shown, and at this price point that's the norm.
You can check every one of our own claims against your own report. That's the point.