AI visibility score
Definition
An AI visibility score is the percentage of unbranded buying questions for which an AI engine names a brand in its answer, sampled over repeated runs because answers vary between identical queries. Three choices decide whether the score is trustworthy: the question set, the sampling method, and the parsing method that decides what counts as a mention.
The three choices that make or break the score
The first choice is the question set. Honest scores are built on unbranded, buyer-intent questions, meaning questions that describe a need without naming any vendor. The reason is mechanical: a question containing the brand's name produces a mention of that brand by construction, so every branded question mixed into the set inflates the score without measuring anything about discovery. The set should also be fixed and versioned, since a score is only comparable over time if the questions underneath it stay constant.
The second choice is sampling. Generative answers are probabilistic, and the variance is larger than intuition suggests: SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same list of recommended brands. A score computed from one run per question is therefore a photograph of noise. Credible scoring runs each question multiple times across days, aggregates into a rate, and treats week-over-week movement as signal only when it exceeds what variance alone would produce.
The third choice is parsing. The rule that keeps a score auditable is evidence-grounded counting: a mention counts only when the brand name literally appears in the stored answer text, so every data point can be traced to the exact words an engine produced. Looser approaches, such as letting a judging model decide whether an answer implied the brand, insert unverifiable interpretation into the score. Reachroller computes its score under exactly these three disciplines, with branded questions excluded from the headline number and every counted mention linked to its raw answer.
How to read and use a visibility score
A visibility score is a rate, and its most useful form is a trend line per engine. The absolute value on any given week matters less than its direction over a month, because absolute values are shaped by category dynamics a brand does not control, such as how list-like the engine's answers are in that category. A score moving from 12 percent toward 30 percent on the engine a brand's buyers actually use is a business result; a static score compared against a vendor's global average is trivia.
Scores should stay disaggregated in two directions. Per engine, because ChatGPT, Gemini and Perplexity retrieve from different sources and a collapse on one can hide inside an average with strength on another. Per question, because the score's whole operational value is the list underneath it: the specific questions where the brand is absent and competitors are named are the ranked backlog of content and third-party work. A score without its question-level detail is a mood rather than a plan.
The commercial context sets the stakes for the number. G2 found 51 percent of software buyers starting research with AI chatbots more often than Google, and 33 percent buying from a brand they had never heard of before an AI named it. A visibility score is the closest available proxy for how often a brand gets to be that named option, which is why it is becoming a standing KPI beside organic traffic rather than a novelty metric.
How visibility scores go wrong
The common failures map onto the three choices. Branded contamination is the most frequent: question sets that include vendor-name queries, sometimes silently, reporting discovery that is actually recognition. Thin sampling is the second, where one run per question turns variance into narrative, and dashboards explain movements that are indistinguishable from noise. Loose parsing is the third, crediting fuzzy matches and implied references that cannot be audited afterward.
There is also a subtler failure: measuring engines through scraped consumer interfaces rather than official APIs, which produces answers shaped by scraping infrastructure, session state and geography in ways that cannot be controlled or reproduced. Reproducibility is the standard to hold any score to. If a vendor or an internal tool cannot show the raw answers behind a number, cannot state its sampling schedule, and cannot list its question set, the score it reports is marketing. The fix is procedural rather than clever: fixed unbranded questions, repeated runs through official interfaces, literal matching, stored answers, per-engine reporting.
One more failure deserves naming because it wastes the most effort: scoring without a work list. A team that receives a number every week but never sees which questions it lost, to whom, and which sources the engine cited, has purchased a mood ring. The score's purpose is to rank the next actions, which means the question-level and citation-level detail underneath it is the product and the headline percentage is its summary. Any scoring setup should be judged by whether a content plan falls out of it directly.
Frequently asked questions
What is a good AI visibility score?+
There is no universal benchmark, because scores depend on how list-like answers are in each category and on the question set used. The useful comparisons are internal: your trend over time per engine, and your rate against the competitors named in the same stored answers. A rising score on the engines your buyers use is the meaningful result.
Why does my score change when I have changed nothing?+
Because generative answers vary between identical runs. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. Sampled thinly, that variance masquerades as movement. A properly sampled score absorbs run-to-run noise into a rate, which is why repeated runs and trend lines are non-negotiable parts of the method.
Should branded questions be part of the score?+
They should be tracked but kept out of the headline number. A branded question guarantees a mention by construction, so including it inflates the score without measuring discovery. Branded tracking earns its keep separately, auditing whether engines describe the brand accurately once its name is in play.
Can one score cover all AI engines?+
A blended score hides more than it reveals. Engines retrieve from different indexes, cite different sources, and move independently, so strength on one can mask collapse on another inside an average. Keep a score per engine and read each engine's trend line separately, weighting them by which engines your buyers actually use.
Sources referenced
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- G2, B2B buyer AI research, 2026
See this metric on your own brand
Reachroller tracks the questions your buyers ask and shows exactly what AI answers. Three days free, no card.
Check my brand free