Concepts

AI share of voice: how to measure your slice of the answer

Updated July 21, 2026

AI share of voice is the percentage of AI assistant answers that mention your brand across a fixed set of buyer questions, benchmarked against the competitors named in the same answers. Computed honestly, it requires four things: an unbranded question set, repeated runs instead of single checks, a separate score per engine, and stored answer text you can audit. Most inflated scores fail on the first requirement, because a question that contains your brand name produces a mention by construction. Reachroller was built around exactly that discipline: branded questions are excluded from the headline score, and every number links to the raw answer behind it. This guide covers the full method, a worked example, and the traps that make the metric lie.

A metric borrowed from PR, and why it fits

Share of voice is one of the oldest measurements in marketing. In its original PR and advertising form, it asked a simple question: of all the conversation in your category, how much of it is about you? Media teams counted press clippings, ad buyers counted impressions against competitor spend, and social teams later counted mentions across platforms. The metric survived every one of those channel shifts because the underlying question never stops mattering. A category has a finite amount of attention, and your slice of it is a leading indicator of your slice of revenue.

AI answers are the newest place that conversation happens, and they concentrate it brutally. When a buyer asks ChatGPT or Perplexity which tools to consider, the answer typically names a handful of brands and silently omits everyone else. There is no page two. There is no position eleven where you still collect a few clicks. You are in the answer or you are invisible, which makes share of voice a more natural fit for AI answers than it ever was for ranked search results.

The buying data explains the urgency. G2 research found that 51 percent of B2B software buyers now start research with an AI chatbot more often than with Google, up from 36 percent just seven months earlier, and that comparing vendor strengths and weaknesses is the single most common AI use case in software research at 41 percent. Forrester's 2026 Buyers' Journey Survey of 18,000 buyers found 94 percent used AI during their most recent purchase, and 55 percent compared vendors inside AI tools. Most striking of all: G2 found 69 percent of buyers chose a different vendor than they originally expected because of AI chatbot output, and 33 percent bought from a brand they had never heard of before an AI named it. Those percentages describe a shelf. AI share of voice measures whether you are on it.

Why the AI version cannot be a single number

Here is where the metric has to be rebuilt rather than copied. Press clippings hold still. A magazine either printed your name in March or it did not, and counting it twice gives the same result. AI answers do nothing of the sort. Large language models sample from probability distributions, which means the same question, asked twice, routinely produces two different answers naming two different sets of brands.

The scale of that instability is measured. SparkToro tested repeated identical prompts and found under a 1 percent chance that two ChatGPT runs return the same brand list. Read that again before trusting any screenshot: the odds that a single run represents what buyers actually see are close to zero. A consultant who checks your brand once and reports a score is reporting a coin flip with a confidence interval wide enough to drive a truck through. We wrote a full breakdown of this problem in why AI gives a different answer every time you ask.

The consequence for share of voice is structural. The metric only becomes meaningful when it is computed over repeated runs, so that the randomness of any single answer averages out into a rate. Appearing in 6 of 20 runs of the same question is a real, stable signal. Appearing in 1 of 1 is an anecdote. Every design decision that follows in this guide flows from that single fact about how these systems generate text.

The honest formula, step by step

Step one: fix the question set.Write 20 to 50 questions your buyers actually ask at each stage: discovery questions such as "what are the best tools for X", comparison questions such as "how does category leader A compare to alternatives", and use-case questions such as "what should a small team use to do Y". Every question must be unbranded with respect to you. The set is fixed before the first run and stays fixed, because a question set that drifts with the results is a score you can no longer compare month over month.

Step two: run repeatedly, per engine. Ask every question on a schedule, through official APIs where they exist, and store the full answer text every time. The stored text is the audit trail. Without it, a share of voice score is an assertion; with it, anyone can open a question and read exactly what the assistant said.

Step three: count mentions strictly.A mention counts only when the brand name literally appears in the stored answer text. No fuzzy matching, no "the assistant probably meant us". This is the rule Reachroller enforces in its own scoring, and it exists because every relaxation of it flatters the number in ways you cannot later defend. The same strict rule is applied to every competitor you track, so the comparison stays symmetric.

Step four: compute two ratios. Answer share is the percentage of answers that mention you at least once: present in 18 of 100 stored answers means 18 percent. Mention share is your mentions divided by all tracked brand mentions: if the same 100 answers contain 120 brand namings and 18 are yours, your mention share is 15 percent. Answer share tells you how often a buyer hears about you at all. Mention share tells you how crowded the answers are and who owns the room. A worked month for a five-brand category might look like this: 25 questions, 4 runs each on one engine, 100 stored answers, and a mention table you can total in a spreadsheet in ten minutes.

A worked example, end to end

Numbers make the method concrete, so here is a full month for a fictional invoicing tool called Ledgerly, tracked against three named rivals. The team writes 25 unbranded questions: eight discovery questions, seven comparison questions, six use-case questions and four budget questions. They run each question four times on ChatGPT across the month, spaced a week apart, storing every answer. That is 100 stored answers for one engine, and with one credit per answer it is also exactly what 100 credits buys in Reachroller's metering, which makes the cost of honest measurement easy to reason about.

Counting strictly, Ledgerly's name appears in 14 of the 100 answers, so answer share is 14 percent. Across all 100 answers the four tracked brands collect 96 total namings: rival A takes 38, rival B takes 27, Ledgerly takes 14, rival C takes 17. Mention share is 14 of 96, or about 15 percent, against rival A's 40 percent. The two ratios agree this month, and they will not always: when answers start naming Ledgerly twice, once in the list and once in a verdict sentence, mention share rises faster than answer share, which is an early signal the engine is warming to the brand.

The real product of the month is the loss list. Sorting questions by Ledgerly's per-question rate surfaces six questions where the brand appeared in zero of four runs while rival A appeared in all four. Those six questions, with the stored answers showing exactly which sources the engine leaned on, are next month's content plan. The score is the headline; the loss list is the work.

The branded-question trap

The fastest way to fake a share of voice score is also the most common: put the brand name in the questions. Ask an assistant "is Acme good for small teams?" and the answer will discuss Acme, because that is what the question demanded. Mix ten of those into a fifty-question set and the headline score jumps by twenty points without a single buyer being more likely to discover you.

This is not a hypothetical failure. Vendor dashboards in this category routinely blend branded and unbranded prompts into one number, and the sales demo looks better for it. When you evaluate any tool, or any agency report, the first question to ask is how branded questions are handled. If the answer is a blank look, the score is inflated. Reachroller's headline score excludes branded questions entirely, and its methodology page documents the exclusion, because a score you cannot explain is a score you cannot defend in a budget meeting.

Branded questions still have a job. They reveal what assistants say about you when asked directly: your pricing, your positioning, whether the description is even accurate. That is sentiment and accuracy monitoring, and it belongs in a separate bucket from discovery measurement. The full argument, with example question pairs, is in branded vs unbranded prompts.

One engine, one score

A blended cross-engine score feels tidy and hides exactly the information you need. The engines see meaningfully different webs. Cross-platform citation analyses, including work published by Profound, have found that only around 11 percent of domains are cited by both ChatGPT and Perplexity. The overlap is that thin because their retrieval stacks differ: 5W Research found Wikipedia and Reddit together account for over a quarter of ChatGPT's U.S. citations, while Perplexity leans even harder on community content and averages about 8.2 sources per answer, roughly 3.4 times ChatGPT's citation density.

Practically, that means a brand can hold a 30 percent answer share on Perplexity and a 4 percent share on ChatGPT at the same time, and a blended 17 percent tells you nothing actionable. Scored separately, the same data tells you precisely where the work is. Which engines deserve your attention first depends on your buyers; the tradeoffs are mapped in which AI engines actually matter for your brand.

There is also a collection-method question hiding here. Scores are only comparable across engines when the answers are gathered the same way, which is why Reachroller tracks ChatGPT, Claude, Gemini, Perplexity and Grok through official APIs only, with ChatGPT live today and the other engines built and rolling out. Scraped consumer interfaces break silently and quietly change what is being measured; an API tells you exactly what it is.

Honest measurement versus inflated measurement

Every share of voice number is downstream of six design decisions. Here is each one, with the version that inflates the score and the version that survives an audit.

Design decisionInflated versionHonest version
Question setQuestions that include your brand nameUnbranded buyer questions only
Run countOne run, screenshottedRepeated runs on a schedule, trended
Mention ruleFuzzy matching, paraphrase countsBrand name literally present in stored text
Engine handlingOne blended score across enginesOne score per engine, compared separately
EvidenceA number with no answers behind itEvery count links to the raw answer
Competitor setChosen after the results are inNamed before the first run

Any one inflated choice can move a score by double digits; stacked together they can manufacture leadership out of absence.

What a good score actually looks like

Calibrate your expectations against the market, because the floor is lower than most teams assume. G2 research found that 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini. Half the market is absent entirely. If your brand shows up consistently on even a handful of unbranded questions, you are already ahead of the median competitor, which is worth remembering before a low first score feels like a crisis.

Resist the urge to chase a universal benchmark. Categories differ in how many brands a typical answer names, engines differ in citation density, and question sets differ in difficulty. The comparisons that mean something are internal and longitudinal: your score this month against your score last month, on the same question set, on the same engine, against the same named competitors. A trend line built that way is one of the few numbers in this discipline that survives scrutiny.

Set the competitor list before the first run and keep it fixed for the same reason the question set stays fixed. Choosing competitors after seeing results invites the quiet habit of dropping whoever is beating you. Three to six named rivals is usually enough to make mention share meaningful without turning the tracking budget into a census. If you want the broader context for why this metric deserves a slot next to your search rankings at all, start with what is AI visibility.

From score to movement

A share of voice score is a diagnosis, and diagnoses do not move markets. The value of the metric is that it hands you a ranked list of losses: specific questions where assistants name rivals and omit you. Each of those is an addressable gap, and the research says the fixes are measurable. The Princeton-led GEO study, published at KDD 2024, found that adding quotations, statistics and cited sources to a page boosted its visibility in generative engine responses by up to roughly 40 percent, while classic keyword stuffing ranked near the bottom of the nine methods tested.

This is the loop Reachroller automates end to end: the tracking produces the ranked losses, and for each lost question it generates a publish-ready fix page with the URL slug, title tag, meta description, schema markup and indexing steps, then a later recheck shows whether the answer flipped. Pricing is metered in credits, where one credit is one AI answer and a generated fix costs ten, with plans from $29 per month. The mechanics are on the how it works page.

If you prefer to run the measurement by hand first, that is a legitimate path, and the full manual protocol, including run counts and storage, is laid out in how to measure AI visibility without lying to yourself. The manual version costs an afternoon a month and teaches you the shape of the problem; the automated version buys the afternoon back and adds the fix loop. Either way, the discipline is the same: fixed questions, repeated runs, strict counting, per-engine scores, stored evidence.

Frequently asked questions

What is AI share of voice?+

AI share of voice is the share of AI assistant answers that mention your brand across a fixed set of buyer questions, measured against competitors. If assistants answer 100 category questions and your brand appears in 18 of them while all tracked brands collect 120 mentions in total, your answer share is 18 percent and your mention share is 15 percent.

How is AI share of voice different from search rankings?+

A ranking is a position on a results page a human still has to click through. AI share of voice measures presence inside the answer itself, where an assistant names a handful of brands and the rest are invisible. There are no page-two consolation slots in a generated answer.

Why should branded questions be excluded from the score?+

Because an answer to a question that contains your brand name mentions you by construction. Counting those answers inflates the score without measuring discovery. Reachroller excludes branded questions from its headline number for this reason and reports them separately.

How many runs do I need before the number is trustworthy?+

More than one, always. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list, so any single run is a coin flip dressed as a metric. Repeated runs on a schedule turn the noise into a trend line you can act on.

Should I combine engines into one blended score?+

No. Cross-platform citation analyses have found only around 11 percent of domains are cited by both ChatGPT and Perplexity, which means the engines see different webs. A blended number hides the engine where you are losing. Score each engine separately and compare.

What is a good AI share of voice?+

There is no universal benchmark, but the floor is low: G2 research found 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini. Any consistent presence on unbranded questions already beats the median. The healthier goal is a rising trend against named competitors, engine by engine.

Sources referenced

  • Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
  • G2, B2B buyer AI research, 2026
  • SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
  • Cross-platform citation overlap analyses of ChatGPT and Perplexity source domains (Profound), 2026
  • 5W Research, ChatGPT citation share analysis, 2026
  • Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)

Get your share of voice measured the honest way.

Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.

Check my brand free