Research

Why AI gives a different answer every time you ask

Updated July 22, 2026

AI assistants give different answers to the same question because they are probabilistic systems, and the effect is larger than most marketers assume. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same list of recommended brands. Every answer is generated fresh through sampling, and when live web retrieval is involved, the pool of sources shifts too. The practical consequence: any single check of your brand's AI visibility is close to meaningless, and any tool that reports one-run scores is selling false precision. The honest measurement is repeated runs over time, with stored answers you can audit, which is exactly how Reachroller computes its visibility scores.

The experiment that put a number on it

The cleanest evidence on AI answer volatility comes from SparkToro, which ran the same brand recommendation prompts through ChatGPT repeatedly and compared the results. The finding: there was less than a 1 percent chance that two identical runs returned the same list of brands. Read that again with a marketer's eyes. If you ask ChatGPT for the best tools in your category right now, and your colleague asks the identical question a minute later, the odds that you both see the same list are close to zero.

This is not a bug OpenAI will patch, and it is not unique to ChatGPT. It is a property of how large language models generate text. Yet almost every conversation about AI visibility still starts with a single screenshot: a founder asks ChatGPT about their category once, sees a rival named, and treats that one sample as the state of the world. The SparkToro result says that a single sample tells you almost nothing about the distribution it came from. The rival in your screenshot might appear in 90 percent of runs or in 4 percent of them, and you cannot tell which from one answer.

The volatility matters commercially because the answers themselves now decide deals. G2's 2026 research found that 51 percent of B2B software buyers start their research with an AI chatbot more often than with Google, and 69 percent chose a different vendor than they originally expected because of what an AI chatbot told them. When an answer that changes with every roll of the dice is steering purchase decisions at that scale, understanding the dice becomes part of the job.

Where the randomness comes from: sampling

A large language model does not look up an answer. It generates one, word by word, and at each step it holds a probability distribution over what could come next. When a model composes a sentence like "the leading options in this category include...", several brand names may sit at comparable probabilities. The decoding process samples from that distribution rather than always taking the single most likely token, which is a deliberate design choice: pure greedy decoding produces repetitive, degenerate text, so production systems keep some randomness, usually controlled by a parameter called temperature.

The consequence for brands is direct. If your brand sits at a similar probability to four competitors for a given question, each run of that question effectively redraws the shortlist. One run names you first, the next omits you entirely, and neither run is wrong in the model's terms. Both are valid samples from the same underlying distribution. What you actually want to know is the shape of that distribution: how often you appear across many draws, and whether that frequency is rising or falling.

This also explains why volatility differs by question. Ask an assistant what the capital of France is and you get the same answer every time, because one token dominates the distribution completely. Ask for the best project management tool for a small agency and you get churn, because dozens of plausible brands crowd the probability space. Commercial recommendation questions, the exact ones that drive revenue, live in the crowded zone. That is worth sitting with: the questions where AI visibility matters most are structurally the questions where answers vary most.

Retrieval adds a second roll of the dice

Sampling would be enough on its own to produce the SparkToro result, but modern assistants add a second source of variance: live web retrieval. When ChatGPT Search, Perplexity or Gemini answers a commercial question, it typically fetches pages from a search index first and grounds the answer in what it retrieved. That retrieval step is itself unstable. Indexes update continuously, freshness signals shift which pages qualify, and the query the assistant silently rewrites from your question can vary between runs.

The engines also differ wildly in how much they retrieve. Perplexity averages about 8.2 sources per answer, roughly 3.4 times what ChatGPT uses, according to citation analyses. More sources means more slots for a brand to enter the answer, but also more ways for the mix to shift between runs. And the source pools barely overlap between engines: cross-platform citation analyses, including work by Profound, found that only around 11 percent of domains are cited by both ChatGPT and Perplexity. An answer is a sample from a source pool, and each engine draws from a different pool.

There is a third, slower layer of change on top: the models and the scaffolding around them get replaced. Providers ship new model versions, adjust system prompts, and rewire retrieval pipelines without notice or changelogs that marketers would ever see. A brand that held steady in answers for months can move overnight because the plumbing changed, and nothing on your side caused it. The distinction between what a model knows from training and what it fetches live matters enormously for how fast you can influence answers, and we cover it in depth in training data vs live retrieval.

Phrasing, context and memory move answers too

Even holding the model and the moment fixed, the way a question is asked shifts the answer. "Best CRM for a startup" and "which CRM should a five-person startup buy" look interchangeable to a human, but they can pull different brand sets because they activate different patterns in the model and different retrieval queries. Buyers do not standardize their phrasing, so a brand that surfaces for one wording and misses another is invisible to a real slice of its market. Any serious measurement panel needs several phrasings per underlying intent, which is why Reachroller tracks a panel of buyer questions per workspace rather than one canonical prompt per topic.

Session context adds another layer. Assistants condition on the whole conversation, so a question asked after ten turns about budget constraints produces a different shortlist than the same question asked cold. Consumer memory features go further: ChatGPT can carry facts about the user across conversations, which means two different people asking the identical question in their own accounts may get systematically different answers. Personalization of this kind is growing, and it quietly breaks the idea that there is one canonical answer to check at all.

For measurement, the implication is clean: audits should run through stateless API calls, where no prior conversation and no user memory contaminates the sample. A logged-in browser session is the worst possible instrument, because it measures the assistant's model of you rather than its model of the market. This is one reason Reachroller queries official engine APIs exclusively rather than scraping consumer interfaces: beyond reliability, API calls are clean-room samples.

The sources of variance, mapped to responses

Each layer of volatility has a measurement response. None of them can be switched off, but all of them can be averaged over or made visible.

Source of varianceWhat happensMeasurement response
Sampling during generationThe model picks each word probabilistically, so two identical prompts divergeRun every question many times and score the frequency of mentions, never a single run
Live retrievalWeb search pulls a different set of pages depending on timing and index stateStore the cited sources per run and watch which domains recur
Model updatesProviders ship new model versions and system prompts without noticeTrack trend lines across weeks so a step change is visible as a step change
Prompt phrasingSmall wording changes shift which brands surfaceTrack a fixed panel of question phrasings, not one canonical wording
Session context and memoryPrior conversation turns and user memory features steer recommendationsMeasure with clean, stateless API calls rather than a logged-in consumer session

The common thread: treat every answer as a sample, never as the truth.

Why one-off checks mislead, with real money attached

The screenshot audit fails in both directions. The optimistic failure: you ask once, see your brand named, and conclude the channel is handled. If that mention was a 20 percent event, four out of five buyers asking the same question never see you, and you have just deprioritized a channel where you are mostly losing. The pessimistic failure: you ask once, see a competitor, and panic into a quarter of reactive content work aimed at an answer that was never representative. Both failures come from treating a random draw as a census.

The stakes behind those failures keep growing. Forrester's 2026 Buyers' Journey Survey of 18,000 buyers found that 94 percent used AI during their most recent purchase, and 55 percent compared vendors inside AI tools. Every one of those comparisons is a fresh sample from the distribution. A brand that appears in 70 percent of runs wins most of those moments; a brand at 15 percent loses most of them. The single number that matters is the frequency, and no single run reveals it.

Volatility also poisons naive before-and-after tests. Publish a new page, ask the assistant once, see your brand appear, and it is tempting to declare the fix worked. But you had some probability of appearing before the fix too. The honest test is frequency before versus frequency after, across enough runs on both sides that the difference exceeds the noise. This is the standard we apply to Reachroller's own fix loop: a fix page ships, indexing steps follow, and a later recheck runs the question again to see whether the answer actually flipped, with the raw answers stored on both sides of the change.

What honest measurement looks like

Once you accept that answers are samples, the measurement method writes itself, and it looks a lot like polling. First, define a fixed panel of buyer questions for your category, covering the intents and phrasings real buyers use. Second, run the panel repeatedly on a schedule through clean API calls. Third, score mentions strictly: a brand counts as mentioned only when its name literally appears in the stored answer text, because fuzzy matching and sentiment inference reintroduce exactly the subjectivity you are trying to remove. Fourth, report the trend, since the direction of your mention frequency over weeks is the signal, and any single reading is noise.

One more correction matters as much as repetition: excluding branded questions. If the question already contains your brand name, the answer will mention you by construction, and counting those answers inflates the score without measuring anything real. Honest AI share of voice is computed on unbranded questions only. We wrote up the full trap in branded vs unbranded prompts, and the broader metric in AI share of voice.

This method is buildable by hand: a spreadsheet, an API key and discipline will get you a defensible baseline, and our guide to measuring AI visibility walks through it step by step. Reachroller automates the whole loop: repeated runs across a question panel, evidence-grounded scoring where every number links to the raw stored answer, branded-question exclusion by default, and trend lines instead of single readings. The methodology is documented openly on the methodology page, because a score you cannot audit is a score you should not trust.

Volatility is also the opportunity

Here is the reframe that makes the randomness bearable: a volatile answer is a movable answer. If AI responses were frozen, a brand missing from them would be locked out until the next model generation. Because answers are redrawn constantly, and increasingly grounded in live retrieval, the composition of the shortlist responds to changes in the underlying evidence. Publish a genuinely citable page, get it indexed, earn a couple of third-party mentions, and your probability of inclusion moves. You are not trying to flip one frozen answer. You are trying to shift a distribution, and distributions shift.

There is measured evidence for how much they shift. The Princeton and Georgia Tech GEO research, published at KDD 2024, found that content optimizations such as adding quotations, statistics and cited sources boosted visibility in generative engine responses by up to 40 percent, with the best methods improving around 22 percent on position-adjusted word count versus baseline. The same study found keyword stuffing performed near the bottom, so the levers that move the distribution are quality levers, and they are the levers you would want to pull anyway.

The volatility even sets the tempo of the work. Because every improvement takes time to propagate through indexes and retrieval, and because verification requires repeated sampling, AI visibility rewards steady weekly cadence over heroic one-time pushes. Track the panel, fix the worst losing question, verify the flip, repeat. Teams that internalize the probabilistic nature of the channel stop chasing screenshots and start compounding. For the broader numbers on how fast this channel is growing underneath the noise, see our AI search statistics roundup.

Frequently asked questions

Why does ChatGPT give different answers to the same question?+

Because large language models generate text by sampling from a probability distribution over possible next words. Unless the temperature is set to zero, two identical prompts can take different paths through that distribution and land on different brand lists. Live web retrieval adds a second layer of variance, since the pages fetched for an answer change with timing and index updates.

How much do AI answers actually vary?+

SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same list of recommended brands. Individual brands often persist across runs, but the exact composition and ordering of the list is different almost every time.

Does answer volatility mean AI visibility cannot be measured?+

No. It means single measurements are unreliable, the same way one poll respondent tells you little but a thousand tell you a lot. Repeated runs of the same question panel produce a stable mention frequency, and trend lines over weeks show real movement. That is the method Reachroller uses.

Is a screenshot of ChatGPT recommending my competitor evidence of a problem?+

It is evidence of one sample, which could be a 90 percent pattern or a 5 percent fluke. Before reacting, run the same question ten or twenty times and see how often the competitor appears. If it recurs in most runs, you have a real gap worth fixing.

Do AI answers vary more for some questions than others?+

Yes. Questions with one dominant, well-documented answer are stable, while open recommendation questions in crowded categories vary the most, because many brands sit at similar probability and sampling decides who makes the cut in each run. Commercial questions, the ones marketers care about, tend to sit in the volatile zone.

Can I reduce the volatility of AI answers about my brand?+

You cannot control the sampling, but you can raise your brand's probability of being included so it survives more runs. The Princeton GEO research found that citable content with statistics, quotations and cited sources improved visibility in generative engine responses by up to 40 percent. Consistent third-party mentions raise the floor too.

How does Reachroller handle answer volatility?+

Reachroller runs your question panel repeatedly through official engine APIs, counts a mention only when the brand name literally appears in the stored answer text, excludes branded questions from the headline score, and reports trends over time. Every number links to the raw answers behind it, so you can audit any score against the evidence.

Sources referenced

  • SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
  • Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
  • G2, B2B buyer AI research, 2026
  • Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
  • Otterly.AI, The AI Citations Report 2026 (1M+ data points)
  • Profound and other cross-platform citation analyses on domain overlap between engines

Stop guessing from screenshots. Measure the distribution.

Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.

Check my brand free