Concepts

Branded vs unbranded prompts: the score inflation nobody mentions

Updated July 20, 2026

A branded prompt contains your brand name; an unbranded prompt describes a need without naming you. The distinction decides whether an AI visibility score means anything, because an assistant answering a question like whether Acme is good for invoicing will mention Acme by construction, while only an unbranded question such as the best invoicing tool for freelancers tests whether the assistant recommends you unprompted. Mixing the two into one number inflates it without measuring discovery, and much of the industry mixes them silently. Honest measurement makes unbranded questions the headline score and tracks branded questions separately for accuracy and sentiment, which is exactly how Reachroller computes its score. This guide covers the mechanics of the inflation, the data on unbranded discovery, and how to build both question sets.

Two kinds of questions, two different measurements

Every question a buyer types into an AI assistant falls on one side of a simple line. Either the question already names a brand, or it describes a need and leaves the naming to the assistant. "Is Notion good for meeting notes?" sits on one side. "What is the best tool for meeting notes?" sits on the other. The words are nearly identical. The measurements they produce are not comparable at all.

A branded question is a request for a description. The assistant retrieves or recalls what it knows about the named brand and summarizes it. The output tells you how the assistant characterizes you: whether your pricing is right, whether the feature list is current, whether the tone is warm or wary. That is genuinely useful information, and it is measurement of a completely different property than visibility.

An unbranded question is a request for a recommendation. The assistant has to choose which brands to surface from everything it knows, and most brands do not make the cut. This is the moment that decides discovery, because it reproduces what a new buyer with no prior brand awareness actually experiences. If your name appears in the answer to an unbranded question, an assistant put it there unprompted. That is the event worth measuring, and it is the only event that belongs in a headline visibility score. The broader case for treating this as a first-class marketing metric is made in what is AI visibility.

Mention by construction: how the inflation works

The arithmetic of the inflation is worth walking through, because it is more violent than it first appears. Suppose a tool tracks 50 questions for your brand and reports the share of answers that mention you. Suppose your true unbranded presence is weak: assistants name you in roughly 10 percent of answers to unbranded category questions. On a clean 50-question unbranded set, your score reads about 10 percent, which is an honest and useful alarm.

Now let 15 of those 50 questions contain your brand name, which is a mix we have seen in real tracking setups. Branded questions produce a mention nearly every time, because the answer is obliged to discuss the thing the question asked about. Fifteen near-guaranteed mentions plus three or four earned ones takes the blended score to roughly 37 percent. Nothing about your market position changed. No new buyer became more likely to hear your name. The number nearly quadrupled anyway, and a dashboard now says things are fine while rivals collect the actual recommendations.

The inflation also compounds over time in a way that hides progress and decay alike. Movement in a blended score is dominated by its noisy unbranded minority, but the branded floor props the level up, so a collapse from 10 percent to 2 percent unbranded presence reads as a gentle dip from 37 to 32. The metric fails in both directions: it overstates health and understates change. Once you see the mechanism, unlabeled blended scores stop being trustworthy anywhere you meet them, in tool dashboards, agency decks and case studies alike.

The data: discovery happens on unbranded questions

The reason unbranded presence deserves the headline slot is that unbranded questions are where buying decisions actually move. G2 research on B2B software buyers found that 51 percent now start research with an AI chatbot more often than with Google, up from 36 percent seven months earlier, and that the single most common use of AI in software research is comparing vendor strengths and weaknesses, at 41 percent. Buyers open a chat window and ask category questions. The brands in those answers form the shortlist.

The consequences show up in vendor selection. G2 found 69 percent of B2B buyers chose a different vendor than they originally expected because of AI chatbot output, and 33 percent bought from a brand they had never heard of before an AI named it. That second number is unbranded discovery in its purest form: a third of buyers ended up purchasing from a company that entered the process through an assistant's unprompted recommendation. No branded prompt can measure your odds of being that company. Only unbranded ones can.

Forrester's 2026 Buyers' Journey Survey of 18,000 buyers rounds out the picture: 94 percent used AI during their most recent purchase, 55 percent compared vendors inside AI tools, and 47 percent built internal business cases before ever contacting a vendor. By the time your sales team hears about a deal, the assistant-mediated shortlisting is largely done. Measuring branded prompts and calling it visibility is measuring the wrong end of that funnel.

The engine mix matters for the question set too. G2 found 72 percent of B2B software buyers use ChatGPT during vendor evaluation, making it the dominant chatbot for software research at 63 percent, while 44 percent also use Perplexity during shortlisting. McKinsey's October 2025 research adds the consumer side: 50 percent of consumers already use AI-powered search intentionally as a primary way to find information and make buying decisions. The same unbranded question can produce very different brand lists on different engines, so run the set on each engine your buyers use and keep the scores separate rather than blending a strong engine and a weak one into a meaningless average.

The same intent, asked both ways

The cleanest way to internalize the split is to see the same buyer intent phrased both ways. Each row below is one intent; only the unbranded column belongs in a visibility score.

Buyer intentBranded promptUnbranded prompt
Category discoveryIs Acme a good project management tool?What are the best project management tools for a startup?
Vendor comparisonHow does Acme compare to its competitors?Compare the leading project management tools on price and features
Use-case fitCan Acme handle a remote team of 50?What should a remote 50-person team use to manage projects?
Budget shortlistIs Acme worth $30 a month?What is the best project management tool under $30 a month?
SwitchingShould I switch from BigCo to Acme?What are good alternatives to BigCo for project management?

One subtlety: a question naming only a competitor, such as alternatives to BigCo, is unbranded for you and branded for BigCo. The label depends on whose score is being computed.

The gray zone: prompts that are hard to classify

Real question sets contain prompts that resist the clean binary, and how you classify them decides whether your score stays honest at the edges. The most common case is the competitor-branded prompt: "alternatives to BigCo" names a brand, but the brand is not yours. For your score, treat it as unbranded, because your appearance in the answer is entirely earned. Just note that these prompts skew easier than pure discovery questions, since the assistant is explicitly fishing for a list of rivals, so report them as their own intent bucket rather than letting them pad the discovery rate.

Comparison prompts that name you and a rival, such as "Acme versus BigCo", are branded for both parties and belong in the branded bucket for both. They measure framing in head-to-head matchups, which is valuable precisely because comparing vendor strengths and weaknesses is the top AI use case in software research according to G2, at 41 percent. Winning the verdict sentence in those answers matters commercially; it just is not discovery, and it must not be scored as if it were.

Then there are the awkward lexical cases. Brands named after common words, a company called Notion or Monday, can appear in prompts coincidentally, and product names can differ from company names, so decide upfront which strings count as your brand and apply the list symmetrically to competitors. Misspellings deserve a policy too: buyers type them, assistants usually resolve them, and a strict literal-match rule should at least document whether common variants are on the match list. None of these decisions is exotic, but each one made silently is a place where a score quietly stops meaning what its audience assumes it means.

Where branded prompts still earn their keep

None of this means branded prompts are worthless. They are the right instrument for a different job: auditing what assistants say when buyers ask about you directly, which late-stage buyers absolutely do. A prospect who has your name from a colleague or an ad will ask an assistant whether you are any good, what you cost, and how you compare to the rival they are also considering. The answer to that branded question can win or kill the deal, and you should know what it says.

Branded monitoring catches three failure modes worth catching. First, factual drift: assistants describing pricing you changed two years ago or features you removed. Second, identity confusion: your brand blended with a similarly named company in another industry, inheriting their reviews and their controversies. Third, sentiment framing: technically accurate answers wrapped in hedging that quietly steers buyers elsewhere. Each failure mode has its own correction path, and we walk through all of them in when AI gets your brand wrong.

The discipline is in the bookkeeping: run branded prompts on their own schedule, report them on their own panel, and never let a branded mention add a point to the discovery score. Two instruments, two dashboards. A thermometer and a scale are both useful, and averaging their readings tells you nothing about either your temperature or your weight.

How the industry handles the split, and how to read scores skeptically

Tooling in this category varies widely on the branded question, and much of it is silent. Free one-shot graders necessarily query the engines about your brand by name, which makes them branded checks by construction: fine as a snapshot of how you are described, and structurally unable to tell you whether anyone discovers you. Subscription dashboards typically let you track any prompts you like, which means the honesty of the score depends entirely on the question set you or the vendor writes, and few vendors document how they label or weight the two types.

Reachroller's position is that the split belongs in the product, and not in the fine print. Branded questions are excluded from the headline score, a mention counts only when the brand name literally appears in the stored answer text, and every number links to the raw answer so the exclusion is checkable rather than claimed. The full scoring rules are public on the methodology page.

Whatever tool or agency you evaluate, three questions expose the handling in under a minute. How many of the tracked prompts contain the brand name? Are branded and unbranded results reported separately? Can I read the raw answers behind the score? Clean answers to all three are rarer than they should be, and a vendor who cannot give them is selling you the inflation described above.

Building your unbranded question set

A good unbranded set starts from buyer language rather than marketing language. Mine your sales calls, support tickets, community threads and search query reports for the phrasings real buyers use, then write 20 to 50 questions spread across the intents in the table above: discovery, comparison, use-case fit, budget and switching. Include the awkward, specific questions, such as tools for a two-person team or options that work without a credit card, because specific questions are where smaller brands actually win mentions against incumbents.

Then hold the set fixed and run it repeatedly. AI answers are probabilistic: SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list, so a single pass over your question set produces noise wearing the costume of a score. Repeated runs on a schedule turn each question into a rate, and rates can be trended, compared and defended. The full computation, including the two ratios worth reporting, is covered in AI share of voice, and the complete manual protocol lives in how to measure AI visibility without lying to yourself.

Expect the honest first number to sting. Unbranded scores for young brands are routinely near zero, and that is the point of measuring: G2 found 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini, so a low score puts you in the majority and hands you a precise list of questions to win. An inflated score would have hidden the same list behind a comfortable number.

What to do with each score once you have it

The unbranded score drives content work. Every unbranded question where assistants name rivals and omit you is a concrete gap, and the research on closing such gaps is unusually clear. The Princeton-led GEO study, published at KDD 2024, tested nine optimization methods and found that adding quotations, statistics and cited sources boosted a page's visibility in generative engine responses by up to roughly 40 percent, while keyword stuffing landed near the bottom. Publish a citable, evidence-dense page that answers the exact question, get it indexed, and recheck the answer after the engines have had a week or two to pick it up.

The branded score drives corrections. Wrong pricing in an answer traces to stale pages or stale third-party content; identity confusion traces to ambiguous entity signals; sour framing often traces to an old controversy or a review-site skew. Each has an escalation path, and none of them is fixed by the content work the unbranded score demands, which is the practical reason the two measurements must stay separate: they feed different work queues.

Reachroller runs both queues from one place: the unbranded tracking produces the ranked losses, each loss generates a publish-ready fix page with slug, title, meta description, schema markup and indexing steps, and a later recheck shows whether the answer flipped. One credit buys one AI answer and a generated fix costs ten, with plans starting at $29 per month. The honest caveat applies: it is a young product, ChatGPT tracking is live today, and the other engines are built and rolling out. But the scoring discipline this article describes is not a roadmap item. It has been the headline number's definition from day one.

Frequently asked questions

What is a branded prompt?+

A branded prompt is any question to an AI assistant that already contains your brand name, such as asking whether your product is good for a task or how it compares to a rival. The answer discusses your brand because the question demanded it, so it measures description quality rather than discovery.

What is an unbranded prompt?+

An unbranded prompt describes a buyer need without naming you: best tools for a job, options under a budget, alternatives to a competitor. It reproduces how new buyers actually discover vendors, so a mention earned on an unbranded prompt is evidence the assistant recommends you on its own.

Why do branded questions inflate AI visibility scores?+

Because the mention is guaranteed by the question itself. If ten of fifty tracked questions contain your name, roughly a fifth of the score is prepaid before any answer is generated. The number rises, but the probability that a new buyer hears your name has not moved at all.

Should I stop tracking branded prompts entirely?+

No, track them in a separate bucket. Branded prompts are how you catch assistants describing your pricing wrong, attributing dead features to you, or confusing you with a similarly named company. They are an accuracy and sentiment instrument, and they should never be blended into a discovery score.

How does Reachroller handle the branded versus unbranded split?+

Reachroller excludes branded questions from the headline score entirely and documents the exclusion in its methodology. A mention counts only when the brand name literally appears in stored answer text from an unbranded question, and every number links to the raw answer behind it.

How many unbranded questions do I need for a meaningful score?+

Twenty to fifty, fixed before the first run, spread across discovery, comparison, use-case and budget intents. Because SparkToro found under a 1 percent chance that two identical ChatGPT runs return the same brand list, each question also needs repeated runs before the rate it produces is trustworthy.

Sources referenced

  • G2, B2B buyer AI research, 2026
  • Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
  • SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
  • Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
  • Vendor pricing and product pages, checked July 2026

See your unbranded score, with the receipts behind it.

Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.

Check my brand free