Concepts

What an AI visibility score actually measures

Updated August 1, 2026

An AI visibility score measures one thing when it is honest: the percentage of unbranded buying questions for which an AI engine names your brand, sampled repeatedly because answers change between runs. The three load-bearing choices are the question set (unbranded, buyer-intent questions only), the sampling method (many runs over time, never one screenshot), and the parsing method (evidence-grounded, so every counted mention links back to a stored raw answer). Get any of the three wrong and the number becomes marketing rather than measurement. Reachroller computes its score exactly this way, excludes branded questions by design, and attaches the raw answer to every data point so you can audit the score instead of trusting it.

A score is a claim about a probability

Start with what the number is supposed to stand for. When a buyer asks an AI engine a question in your category, the engine composes one answer and names a handful of brands. Your visibility score is an estimate of how often that handful includes you. It is a probability dressed as a percentage: given a relevant question and a fresh conversation, what are the odds the answer says your name?

That framing matters because probabilities have measurement rules. You cannot estimate a coin's bias from one flip, you cannot estimate it from flips of a different coin, and you cannot report it honestly if you throw away the record of the flips. Every failure mode in AI visibility scoring is one of those three mistakes wearing a dashboard. The stakes are real: G2 found 51 percent of B2B software buyers now start research with an AI chatbot more often than Google, so the probability this score estimates is the probability of existing in the buyer's first conversation.

The rest of this article walks the pipeline an honest score runs through, then flips it around and catalogs the tricks that inflate a dishonest one. If you are newer to the category, the ground-level primer is what is AI visibility; this piece assumes you know why the answer matters and asks how to measure it without lying to yourself.

Foundation one: unbranded questions only

The question set is the single largest lever on the final number, and the branded question is the largest lever inside it. Ask ChatGPT "what is Acme and is it any good?" and the answer discusses Acme, because the question forced it to. Ask "what is the best tool for X?" and Acme has to earn the mention. These are different measurements, and only the second one predicts whether new buyers discover you.

The 2026 data shows how far apart the two readings sit. Victorious tested brands across eight AI platforms and found 96 percent were described accurately when asked about directly, yet 89 percent never appeared in answers to category research questions. The same brand can score near 100 on branded prompts and near zero on unbranded ones, on the same engine, in the same week. A vendor that averages the two sells you a blend of a solved problem and an unsolved one, and the blend always flatters. The full argument is in branded vs unbranded prompts.

Unbranded is necessary but insufficient: the questions also have to be the ones buyers ask. A set stuffed with obscure long-tail phrasings nobody types will read differently from a set built on the four intents buyers actually bring to assistants, the best-of request, the alternatives request, the head-to-head comparison and the how-to question. Reachroller structures every tracked question set around those intents so the score reflects the conversations that shortlist vendors, and the taxonomy is documented in the four buyer intents in AI prompts.

Foundation two: repeated sampling, because answers move

AI answers are generated fresh each time, and the generation is stochastic. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. Same question, same engine, same minute: different brands, different order, different citations. Any score built on one run per question is therefore a photograph of noise. It can swing ten or twenty points between Monday and Tuesday with no change in the underlying reality.

The statistical fix is boring and non-negotiable: sample each question multiple times, spread the samples across days, and report the rate. If an engine names you in 6 of 20 runs on a question, your presence on that question is 30 percent, and next month's 45 percent is a real improvement rather than a re-rolled die. Rates also make volatility itself visible, which matters because a brand mentioned erratically has a different problem from a brand never mentioned at all. Why the answers move this much is its own subject, covered in why AI gives a different answer every time you ask.

Sampling hygiene has quieter requirements too. Each run needs a fresh session, because a conversation that already mentioned your brand contaminates every later answer in it. Runs should avoid logged-in personalization, which bends answers toward past behavior. And the schedule should be fixed rather than opportunistic, because checking more often when things look good is p-hacking with extra steps.

Foundation three: evidence-grounded mention parsing

Between the raw answer and the score sits a parser, and the parser is where quiet inflation lives. Deciding whether an answer "mentions" a brand sounds trivial and is not. Brand names collide with common words. Ask about project management and a sloppy substring matcher credits "Monday" every time the answer says "on Monday mornings". Brands get mentioned negatively, as in "avoid X for this use case", which some tools happily count as visibility. And answers sometimes name a brand while citing a competitor's comparison page, which is a different event from citing you.

Evidence-grounded parsing means the pipeline stores the full raw answer, detects mentions against that stored text, distinguishes a mention from a citation, and keeps the linkage so every counted event can be opened and read. This is how Reachroller computes its score: each tracked question run produces a stored answer, the mention decision is made on that answer, and the report shows the receipt next to the data point. The distinction between being named and being linked is worth its own read in AI mentions vs citations.

The audit property is the point. A score you can audit converts arguments about methodology into two minutes of reading. When a founder doubts a low number, the fastest resolution is opening five raw answers and seeing the competitors named instead. When an executive doubts an improvement, the receipts show the same question flipping from a loss to a win across dated runs. Measurement you cannot audit is testimony; measurement you can audit is evidence.

The honest pipeline, end to end

StageHonest versionFailure mode
1. Question setUnbranded buying questions a real prospect would askBranded questions mixed in, guaranteeing mentions
2. SamplingRepeated runs on a schedule, fresh sessions each timeOne run, screenshotted on a lucky day
3. Answer captureFull raw answer stored verbatim, with citationsOnly a yes/no flag kept, evidence discarded
4. Mention parsingExact brand detection, checked against the stored textFuzzy substring matching that counts false positives
5. AggregationPlain mention rate across questions and runsOpaque weighting that cannot be recomputed by hand
6. ReportingScore plus per-question receipts and trend lineA single grade with no way to see what produced it

Each stage can be verified independently. If a vendor resists showing any one of them, assume that stage is where the inflation lives.

The inflation tricks, and how to catch each one

The market pressure to inflate is structural. A tool that reports a flattering score gets renewed; a tool that reports an honest zero has to argue for its life. So the tricks below are common, they are rarely framed as tricks, and several of them can hide inside a genuinely pretty product. The defense is knowing what each one does to the number.

The branded-question blend is the heavyweight, capable of moving a score from single digits to the sixties on its own, which is exactly the gap the Victorious data documents between direct and category questions. Single-run snapshots are the most common, because continuous sampling costs real compute and a screenshot costs nothing. Cherry-picked question sets are the subtlest: track twenty questions the brand already wins, retire the ones it loses, and the score climbs while reality stands still. Fuzzy matching inflates mechanically and randomly. And opaque composite grades launder all of the above into a letter that no one can recompute.

TrickWhat it does to the scoreHow to detect it
Branded questions in the setMentions by construction; score jumps 20 to 60 pointsAsk for the question list; count questions naming the brand
Single-run snapshotsA lucky answer becomes the permanent scoreAsk how many runs per question feed the number
Cherry-picked question setsOnly questions the brand already wins get trackedAsk who wrote the questions and when they last changed
Fuzzy or partial matchingGeneric words counted as brand mentionsOpen raw answers and search for the brand string yourself
Opaque composite gradesWeighting hides losses inside a friendly letter gradeAsk for the formula; an honest score recomputes by hand

Two questions expose most vendors in one call: "show me the exact question list" and "show me the raw answer behind this data point".

A worked example: one question, one month, one honest number

Make it concrete. Suppose you sell invoicing software and track the question "what is the best invoicing tool for freelancers?" on ChatGPT. Over the month, the schedule runs the question eight times in fresh sessions. Your brand appears in runs two, five and seven, so your mention rate on this question is 3 of 8, or 37.5 percent. Each of the eight answers is stored verbatim: the three that name you, and the five that name FreshBooks, Wave and a Reddit-thread favorite instead, along with the sources each answer cited. Nothing about the number requires trust; anyone can open the eight receipts and recount.

Now scale the arithmetic. Twenty-five tracked questions at eight runs each is 200 stored answers. If 58 of them mention you, the headline score is 29 percent, and the per-question table shows exactly which conversations produced the 58 and which produced silence. Watch what one branded question would do to this math: add "is your brand good for freelancers?" and its 8 guaranteed mentions, and the score climbs to roughly 32 percent without a single buyer-facing answer changing. Four free points from one polluted question, and most polluted sets contain more than one.

The worked example doubles as a vendor test. Hand any AI visibility vendor this scenario and ask them to walk you from their raw answers to their score with a calculator. An honest methodology survives the walk in five minutes. A methodology that needs a weighting you cannot see, a proprietary index, or a "confidence adjustment" is asking you to grade its homework without showing the working.

What even an honest score cannot tell you

Honesty includes limits. A visibility score does not measure sentiment: an engine can name you while describing you inaccurately, and fixing that is a correction problem rather than a visibility problem. It does not measure position or emphasis unless the methodology explicitly weights them, and if it does, the weighting should be published. It does not measure revenue, though it correlates with the pipeline metrics that do: Semrush found AI visitors convert at 4.4 times the rate of standard organic, which is why the mention that precedes the visit is worth measuring at all.

A score also cannot tell you why you lost a question, only that you did. The why lives in the citations: McKinsey's August 2025 AI Discovery Survey found a brand's own website supplies only 5 to 10 percent of the sources AI platforms reference, and University of Toronto research put third-party content at 91 percent of cited sources. Diagnosis therefore requires storing which sources the engine used, per lost question, which is a second reason raw answers must be kept. Scores summarize; receipts explain.

Finally, a score is engine-specific. ChatGPT, Gemini, Perplexity, Claude and Grok retrieve differently and cite different source mixes, with one analysis finding only about 11 percent of domains cited by both ChatGPT and Perplexity. A single blended cross-engine number hides the fact that you may be a leader on one engine and invisible on another, so honest reporting keeps per-engine breakdowns available even when it leads with one number.

How Reachroller computes the number it shows you

Reachroller's score is a plain, recomputable mention rate over unbranded buying questions. You track a fixed set of questions built on the four buyer intents. The platform runs them against AI engines on a schedule, each run in a fresh session, stores every raw answer verbatim with its citations, and parses mentions against the stored text. The headline score is mentions divided by sampled answers, branded questions excluded by design, and every data point links to its receipt. The methodology is public on our methodology page, because a score you have to take on faith is the thing this product exists to replace.

The honest caveats apply here too. Reachroller is a young product: ChatGPT tracking is live today, and the other four engines are built and rolling out. The score will sometimes be low, because for most brands the true number is low; Wellows found over 73 percent of brands with page-one Google rankings have zero AI mentions. What you get for the low number is a per-question map of exactly where you lose and a generated, publish-ready fix page for each loss, which is the part of the loop that moves the score. Starter is $29 per month for 400 credits and 25 tracked questions, where one credit is one AI answer and a fix page costs ten.

If you are comparing measurement approaches across the market before committing, the wider tooling landscape, including the enterprise platforms and their methodologies, is mapped in the best AI visibility tools. Whichever you choose, hold it to the three foundations: unbranded questions, repeated sampling, receipts for every mention. A score built on anything less is a mood ring with an API.

Frequently asked questions

What is an AI visibility score?+

It is the percentage of tracked, unbranded buying questions for which an AI engine such as ChatGPT, Claude, Gemini, Perplexity or Grok mentions your brand, measured across repeated runs. A score of 32 means the engine named you in roughly a third of sampled answers to questions where a buyer did not type your name.

Why must branded questions be excluded from the score?+

Because they guarantee a mention. Victorious tested eight AI platforms in 2026 and found 96 percent of brands were described accurately when asked about directly, while 89 percent never appeared in answers to category research questions. A score that mixes both question types averages a solved problem into an unsolved one and reports the blend as progress.

How many runs does a reliable score need?+

More than one, always, because AI answers are probabilistic. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. In practice, honest tools sample each question multiple times per reporting period and present a rate, so a single unusual answer moves the score a little instead of defining it.

What does evidence-grounded parsing mean?+

It means every counted mention traces back to a stored raw answer you can open and read. The parser detects the brand in the actual answer text, and the report links the score to those receipts. If a tool cannot show you the answer behind a data point, the data point is an assertion rather than a measurement.

Is a low AI visibility score bad news?+

It is normal news. Wellows found over 73 percent of brands have zero mentions in AI responses despite ranking on Google page one, and the Victorious study put the invisible share at 89 percent for category questions. A low first score is the honest baseline most brands start from; the useful signal is the trend after you start publishing fixes.

How does Reachroller compute its visibility score?+

Reachroller tracks a fixed set of unbranded buying questions per brand, runs them against AI engines on a schedule, stores every raw answer, parses mentions against that stored text, and reports the mention rate with per-question receipts. Branded questions are excluded from the headline number by design, and the methodology is public. Starter is $29 per month.

Sources referenced

  • Victorious, AI brand mention study across eight AI platforms, Q2 2026 (reported by Search Engine Journal)
  • SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
  • Wellows, GEO visibility research on brands with page-one rankings, 2025
  • McKinsey, AI Discovery Survey on source composition of AI answers, August 2025
  • University of Toronto, analysis of third-party citation share in AI answers, 2026
  • Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
  • Semrush, 2026 AI Visibility Index, 126 million U.S. AI search prompts

Get a score you can audit, not admire

Three days, 50 credits, every feature, no card. Unbranded questions, repeated runs, and the raw answer behind every data point.

Check my brand free