Playbooks
How to measure AI visibility without lying to yourself
Updated July 24, 2026
Measuring AI visibility honestly requires four disciplines that most quick checks skip. Run every question repeatedly, because SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list, so any single reading is noise. Store the raw answer text so every score can be audited. Count a mention only when the brand name literally appears in that stored text. And exclude branded questions from the headline number, because an answer to a question containing your name mentions you by construction. Apply those four rules across a fixed set of buyer questions and you get a trend line you can defend. This is the method Reachroller runs automatically; the full manual protocol is below.
Why this metric is so easy to fake, even to yourself
AI visibility has a measurement problem that classic SEO never had. A Google ranking is stable enough to check once and cite for a month. An AI answer is regenerated on every request, from a probabilistic model, often with live retrieval mixed in, and it changes between two identical asks. SparkToro put a number on it: under a 1 percent chance that two identical ChatGPT runs return the same list of recommended brands. Whatever you saw when you checked this morning, the next buyer probably saw something different.
That volatility makes self-deception nearly frictionless. Ask until the answer flatters you, screenshot it, and you have evidence for the board deck. Run one check after publishing a page, see your name, and conclude the work succeeded, when the same check an hour later would have shown nothing. The deception works in both directions: a founder can panic over one bad answer that the distribution would have smoothed out. The mechanics behind the churn are unpacked in why AI gives a different answer every time you ask.
The stakes for getting this right are not academic. G2's 2026 research found 69 percent of B2B software buyers chose a different vendor than they expected because of AI chatbot output, and 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini. Budgets are moving toward this channel, and budgets follow numbers. If the numbers are theater, the spend is too. What follows are four principles that make the number real, then a protocol for producing it by hand.
Principle 1: repeated runs, scored as a frequency
The unit of honest AI visibility measurement is a frequency across repeated runs, never a single answer. If your brand appears in three of five runs of a question, your visibility on that question is 60 percent this period. That framing absorbs the volatility instead of being fooled by it, and it converts the SparkToro finding from a reason measurement is impossible into a design requirement for doing it properly.
In practice, three to five runs per question, spread over several days rather than fired in one burst, is the workable floor. Spreading matters because engines with live retrieval can shift as their indexes refresh, and you want your sample to average over that churn rather than capture one moment of it. More runs are better and cost linearly more; five runs per question across 25 questions is 125 answers per engine per period, which is where hand-run measurement starts to strain and metered tooling starts to look cheap.
The corollary is that trend lines are the only comparison worth making. A 60 percent reading this month against 20 percent last month, on the same questions with the same protocol, is a real improvement. A 60 percent reading against a rival's single lucky screenshot is not a comparison at all. Freeze the protocol, then watch the line.
Principle 2: store the raw answers
Every score should be one click away from the evidence that produced it. Store the full text of every answer, the favorable ones and the ones that recommended your competitor, with the question, engine, and timestamp attached. This is the difference between a measurement and a claim: a measurement can be audited by a skeptic, including the most important skeptic, which is you in three months trying to remember why the March number looked so good.
Stored answers also carry the diagnostic payload that a bare score throws away. The raw text tells you which competitors appear alongside you, what the engine says about your pricing and features, which sources it cited, and whether a miss was a near-miss or total absence. When a number moves, the stored answers tell you why it moved, which is the difference between a dashboard you glance at and a dataset you act on.
This principle is also the sharpest test to apply to any tool you evaluate. If a product shows you a visibility score but will not show you the raw stored answers behind it, it is asking you to trust arithmetic you cannot check. Reachroller's methodology makes every number link to the raw answer text it was computed from, and any measurement setup, hand-built or bought, should meet that bar.
Principle 3: count only literal mentions
A mention counts when the brand name literally appears in the stored answer text. That rule sounds pedantic until you watch scoring drift without it. Did the engine mention you when it described "tools like yours"? When it linked your domain without naming you? When it named your product but attributed it to the wrong company? Fuzzy matching answers each of these generously, and every generous call inflates the score in a way nobody can verify later.
Literal counting is strict, reproducible and slightly conservative, which is the right bias for a number that will justify spending. It also forces a useful companion distinction: a mention, where your name appears in the answer text, and a citation, where your site is linked as a source, are different events with different mechanics, and mixing them into one number obscures both. Track them as separate columns. The distinction gets a full treatment in mentions vs citations.
One refinement worth adopting from the start: log the context of each literal mention, at minimum whether it was a recommendation, a neutral reference, or a negative comparison. The headline score counts all literal mentions, and the context column tells you whether to celebrate them. An engine that names you as the option to avoid is visibility of a kind, but it belongs in a different conversation.
Principle 4: exclude branded questions from the score
Ask an engine "is Acme good for invoicing" and the answer will discuss Acme, every time, by construction. Fold answers like that into your visibility score and the score rises without meaning anything, because the question did the mentioning, and no buyer who has not already heard of you will ever ask it. This is the single most common inflation in AI visibility reporting, and plenty of tools quietly commit it.
The honest score is computed over unbranded questions only: the category questions, comparisons and problem statements a buyer asks before they know you exist. Those are the answers where being named is an achievement, because the engine chose you from the whole market. Reachroller enforces this split automatically and excludes branded questions from the headline score; if you measure by hand, tag every question as branded or unbranded before the first run, and never let the tags blur. The full argument, with worked examples, is in branded vs unbranded prompts.
Branded questions still deserve tracking, in their own bucket, for a different purpose: accuracy. The answer to a branded question tells you what engines believe about your pricing, features and positioning, and it is where factual errors about your brand surface first. Score visibility on unbranded questions; monitor truth on branded ones.
Putting it together: AI share of voice
With the four principles in place, the summary metric almost assembles itself. AI share of voice is your mention frequency across all unbranded runs in a period, reported per engine, ideally next to the same figure for the competitors you name. Mentioned in 41 of 125 unbranded ChatGPT runs is 33 percent share of voice on that engine this period; a rival at 55 percent tells you who currently owns your category's answers. How to compute and present it, including the traps, is covered in AI share of voice.
Report it per engine rather than blended, because engines disagree more than most people expect. Cross-platform citation analyses find only about 11 percent of domains are cited by both ChatGPT and Perplexity, so a brand can legitimately hold 50 percent share of voice on one engine and near zero on another. A blended average would hide exactly the gap you need to see to prioritize work.
And always publish the denominator with the number. "Mentioned in 12 answers" is theater; "mentioned in 12 of 60 unbranded runs across 20 questions on ChatGPT, up from 5 of 60 last period" is a measurement. Anyone reporting the first form, vendor or colleague, should be asked for the second.
The mistakes, mapped to their fixes
| Common practice | Why it lies | Honest alternative |
|---|---|---|
| One run per question | Answers are probabilistic; under 1% of identical run pairs match (SparkToro) | Three to five runs per question, spread across days, scored as a frequency |
| Screenshot as evidence | A favorable screenshot is one draw from a distribution, chosen after the fact | Store every raw answer, favorable and not, before computing anything |
| Counting branded questions | A question containing your name gets you mentioned by construction | Score unbranded questions only; track branded ones separately for accuracy |
| Fuzzy mention matching | Counting near-matches and implications inflates the score unverifiably | Count a mention only when the brand name literally appears in the answer text |
| Changing the question set | New questions each month make every period incomparable to the last | Freeze a core question set; version any additions explicitly |
| Reporting a score with no denominator | "Mentioned in 12 answers" means nothing without runs and questions counted | Report mentions over total unbranded runs, per engine, per period |
Every row applies equally to hand-run spreadsheets and to tools you are evaluating; the practices are what matter, not the software.
The manual protocol, start to finish
Here is the whole method as a runnable procedure. Write 20 to 25 unbranded buyer questions and tag each with the competitor you suspect owns it. Pick the engines that matter for your market. Run every question three to five times per engine across a few days, pasting each full answer into a log with question, engine, run number and date. Then score: for each answer, record whether your brand name literally appears, which competitors appear, and any citations of your domain. Compute mention frequency per question and share of voice per engine. Date the summary and freeze it as your baseline.
Repeat the identical procedure every two to four weeks, same questions, same run counts, and chart the trend. Expect the first repeat to be humbling: numbers that looked like progress often regress toward the distribution, and that regression is the method working. A fuller walkthrough of the first cycle, including how to choose the questions, is in the AI visibility audit.
Budget honestly for the cost: the first full cycle on one engine is the better part of a day, and the value depends entirely on repeating it on schedule, which is where hand-run measurement usually dies. This is the calculation behind metered tooling. Reachroller runs this exact protocol through official engine APIs, one credit per answer, from $29 per month for 400 credits and 25 tracked questions, with the raw answers stored and the branded exclusion built in. The method is identical; what you are buying is the cadence and the audit trail. Pricing details are on the pricing page.
Connecting visibility to outcomes, carefully
Eventually someone will ask what the visibility number buys. Answer with the conversion evidence, and present it with its spread intact. WebFX, analyzing 2.3 billion sessions across 2024 and 2025, found AI-referred visitors converting at about 1.2 times organic. Semrush's 2026 figure is higher, around 4.4 times standard organic, and Adobe reported in March 2026 that AI traffic converted 42 percent better than non-AI. Opollo's AI Search Benchmark measured 14.2 percent conversion for AI-referred visitors against 2.8 percent for Google organic. Ahrefs found up to 23 times higher conversion, a figure widely treated as an outlier and worth citing only with that caveat attached.
Present the honest denominator alongside: AI referrals are still around 0.18 percent of all sessions in some panels, and much AI-sourced traffic hides in direct because assistants often do not pass referrers. The channel is small and high-intent, growing fast, with ChatGPT commanding roughly 92 percent of trackable LLM referral traffic. SE Ranking's May 2026 study across 101,574 websites found ChatGPT referrals jumping 36.7 percent in a single month to an all-time high. Small, compounding, and disproportionately likely to convert is the accurate summary.
The reporting structure that follows: AI share of voice as the leading indicator you can move this quarter, AI referral traffic and its conversion rate as the confirming indicator that follows with a lag. Both measured with the same honesty, both trending in the same deck. That combination survives any skeptic in the room, which is the entire point of measuring this way.
One last framing that helps with executive audiences: present the conversion spread as a spread. Saying AI-referred visitors convert somewhere between 1.2 and 4.4 times organic depending on the study, with one 23 times outlier excluded, is more credible than quoting whichever single figure flatters the slide, and it inoculates the whole deck against the first person who has read a different study. Teams that report ranges with sources earn the right to be believed about their own numbers, and in a channel this young, that earned trust is itself a competitive asset.
Cadence, targets, and what good progress looks like
Once the protocol is running, two practical questions remain: how often to measure, and what counts as progress. On cadence, every two to four weeks is the useful range for a brand actively publishing fixes. Faster than weekly mostly samples noise, because the interventions that move answers, new pages entering indexes and third-party mentions being crawled, play out on one to two week timelines. Slower than monthly and you lose the ability to connect a specific change to a specific movement, which is half the value of measuring at all.
On targets, anchor expectations to the mechanism rather than to hope. A question you lose because the winning answer cites a live page is winnable within a cycle or two of publishing something better. A question the engine answers from training data can sit unmoved for months no matter how good your new page is, and the honest response is to log it as training-data-bound and revisit quarterly rather than burn cycles rechecking it weekly. Segmenting your question set this way keeps the trend line readable: movement where movement is possible, patience where it is not.
Good progress in practice looks unglamorous: share of voice on winnable unbranded questions climbing a few points per cycle, the branded-question bucket staying factually clean, and the stored answers showing your name appearing in more runs of the same questions rather than in cherry-picked new ones. A quarter of that produces a defensible before-and-after. One more habit protects it: when you change the question set, version it, report old and new sets separately for one overlapping period, and never splice the two lines into one chart. Continuity of method is what makes the chart mean anything.
Frequently asked questions
Why is a single AI visibility check misleading?+
Because AI answers are probabilistic. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. A single check is one draw from a wide distribution: it can show you present when you are usually absent, or absent when you usually appear. Repeated runs scored as a frequency are the minimum honest unit of measurement.
What is a branded question, and why exclude it from the score?+
A branded question contains your brand name, like asking whether your product is good. The answer will mention you by construction, so counting it inflates visibility. Unbranded category questions, the kind buyers ask before they know you exist, are the only fair test of whether AI recommends you. Branded questions are still worth tracking separately to catch factual errors about your brand.
How many questions and runs do I need for a usable baseline?+
A workable floor is 20 to 25 unbranded buyer questions, each run three to five times over a few days, per engine you care about. That is 60 to 125 answers per engine, enough to compute mention frequency with some stability. Fewer questions or single runs produce numbers that swing too wildly to support any decision.
What is AI share of voice?+
The share of AI answers in your category that name your brand, usually alongside the same figure for named competitors. Computed honestly, it is mentions divided by total unbranded runs, with the same repeated-runs discipline underneath. It borrows the concept from PR measurement and is the most decision-useful single number in AI visibility, provided branded questions stay out of it.
Can I tie AI visibility to revenue?+
Partially. AI referral traffic is still small, around 0.18 percent of sessions in some panels, and much of it hides in direct traffic, but it converts unusually well: WebFX measured AI-referred visitors converting about 1.2 times organic across 2.3 billion sessions, Semrush reported roughly 4.4 times standard organic, and Adobe found 42 percent better conversion. Treat visibility as the leading indicator and referral conversion as the confirming one.
Do I need a tool, or can I measure by hand?+
The method works by hand: fixed questions, repeated runs, logged answers, literal counting, branded exclusion. The cost is time, roughly a day per full cycle for 25 questions on one engine, repeated every few weeks. Tools earn their keep on cadence and audit trail. Reachroller runs the same method through official engine APIs from $29 per month, stores every raw answer, and links every score to the evidence behind it.
Sources referenced
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- WebFX, AI traffic growth and conversion analysis, 2.3B sessions, 2024-2025
- Semrush, AI-driven visitor conversion analysis, 2026
- Adobe, AI traffic conversion analysis, March 2026
- Opollo, AI Search Benchmark (AI-referred vs organic conversion)
- Ahrefs, AI referral conversion analysis (cited as an outlier in text)
- SE Ranking, ChatGPT referral traffic study, May 2026 (101,574 websites)
- G2, B2B buyer AI research, 2026
- Profound and cross-platform citation analyses (engine overlap)
Get a visibility score you can actually defend.
Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.
Check my brand free