Engines

How ChatGPT decides which brands to recommend

Updated July 25, 2026

ChatGPT recommends brands through two mechanisms. The first is training data: patterns absorbed from the public web during pretraining, which only change when OpenAI retrains a model. The second is live retrieval: when a question needs current information, ChatGPT searches the web, reads pages, and assembles an answer with citations, drawing on an index that launched on Bing and is increasingly fed by OpenAI's own crawler, OAI-SearchBot. Citation studies show it favors reference and community sources, with Wikipedia and Reddit together supplying over a quarter of its citations. The picks are also probabilistic, so the same question can name different brands on different runs. Reachroller tracks exactly this pipeline: it runs your buyers' questions repeatedly, shows which answers name you, and generates the citable pages that flip the ones that do not.

Two pathways to a recommendation

When a buyer types "what is the best project management tool for a small agency" into ChatGPT, the answer is assembled through one of two pathways, and everything about influencing the result depends on which one fired. The first pathway is parametric knowledge: the statistical patterns the model absorbed from its training corpus. If your brand appeared frequently enough on the public web, in reviews, comparisons, documentation, forum threads and news coverage, the model can name you and describe you without touching the live web at all. This is the pathway that answers instantly, with no citations, and it reflects the web as it looked when the training snapshot was taken.

The second pathway is retrieval. When the model judges that a question needs current information, and buying questions with prices, feature lists and "best in 2026" framing usually do, ChatGPT issues web searches, reads a set of pages, and composes an answer grounded in what it just read, with citation links back to the sources. This is the machinery behind ChatGPT Search, and it has its own supply chain of indexes and crawlers, which we cover in depth in how ChatGPT Search works.

The distinction matters because the two pathways move on completely different timelines. Training data changes when OpenAI retrains a model, on a schedule you cannot see or influence. Retrieval changes when the pages it reads change, which means a brand can alter retrieval-driven answers in weeks by publishing and indexing the right content. If you take one idea from this article, take this one, and if you want the full breakdown, we wrote a dedicated guide to training data versus live retrieval.

The scale behind the answers

The reason this machinery deserves your attention is its audience. ChatGPT reached roughly 900 million weekly active users in early 2026, about double the 400 million it reported in February 2025, and is on pace to cross a billion. Its search experience alone processes an estimated 250 to 500 million queries per week. Those are consumer-scale numbers, and the buying behavior inside them has been measured. G2's 2026 buyer research found that 72 percent of B2B software buyers use ChatGPT during vendor evaluation, making it the dominant chatbot for software research at 63 percent, and that 51 percent of buyers now start research with an AI chatbot more often than with Google, up from 36 percent just seven months earlier.

The traffic side tells the same story. ChatGPT commands about 92 percent of trackable LLM referral traffic, and that traffic grew roughly 12.8 times over 19 months. SE Ranking's study of 101,574 websites found ChatGPT referrals jumped 36.7 percent in May 2026 alone, an all-time high. The honest caveat, and we always include it: AI referrals remain a small share of total sessions in most panels, around 0.18 percent in some measurements, and a lot of AI-sourced traffic hides inside "direct". The volume is small. The intent is not, and the recommendation happens before the click either way.

Put those two facts together and the stakes become concrete. A brand that ChatGPT consistently names when buyers ask category questions collects shortlist placements at a scale no single publication can match. A brand it never names is invisible at the exact moment 63 percent of software researchers are forming their first list. Forrester's 2026 Buyers' Journey Survey of 18,000 buyers found 94 percent used AI during their most recent purchase, and 55 percent compared vendors inside AI tools. The recommendation engine you are reading about is already sitting in the middle of your funnel.

What the training data pathway rewards

The training pathway rewards breadth and consistency of coverage across independent sources. A model learns to associate a brand with a category the same way it learns everything else: by seeing the association repeated, in varied wording, across many documents. Brands with years of reviews, comparison articles, community discussion and documentation on the open web are heavily represented in the weights. Young brands, brands that market mostly inside closed platforms, and brands whose content lives behind logins are thin or absent. This is why long-established names often dominate no-search answers even in categories where they are no longer the best choice.

Two properties of this pathway deserve respect. First, it is slow in both directions: coverage you earn today shows up in some future model, and mistakes fossilize the same way, which is why models sometimes describe pricing or features a company retired years ago. Second, it is uneven. There is no query log to study and no index to submit to, so the only strategy is the patient one, being written about accurately and often, in places the next training snapshot will include. That is real work with compounding value, but it is not a campaign with a deadline.

For a working brand, the practical conclusion is about sequencing. You cannot schedule a retraining, so treat parametric answers as a lagging indicator, a measure of your accumulated reputation on the open web. Focus active effort on the retrieval pathway, where cause and effect operate on publishing timelines, and let sustained coverage handle the weights. Reachroller's question runs make the split visible in practice: answers that arrive without citations are telling you about the past, and answers with citations are telling you exactly which live pages to influence this month.

The retrieval pathway: from Bing seed to OAI-SearchBot

ChatGPT's search experience launched on Bing's index, and Bing remains part of the supply chain. But OpenAI has been building independence: it operates its own crawler, OAI-SearchBot, and a Botify analysis found OpenAI roughly tripled its web crawl since August 2025. The direction of travel is clear. ChatGPT increasingly reads the web through its own eyes, which means publishers now have a direct relationship with OpenAI's crawler in their server logs, and blocking it in robots.txt has a direct cost in retrieval visibility.

For a brand, the retrieval pathway has a precondition that surprises people: classic indexing still gates everything. A page that search engines cannot find and rank is effectively invisible to retrieval, because the candidate set ChatGPT reads from is assembled through search-style lookups. Submitting new pages through Google Search Console and Bing Webmaster Tools, keeping them crawlable, and giving them internal links remains step one. This is also why every fix page Reachroller generates ships with indexing steps attached: publishing without indexing is shouting into a room the engine never enters.

Once a page is in the candidate set, the question becomes selection: out of everything retrievable for a query, which handful of pages does the model actually read and cite, and which brands inside them make it into the recommendation? That is where the citation studies come in, and they paint a picture most marketing teams have not internalized yet.

What ChatGPT actually cites

The most striking finding in the citation research is how little ChatGPT's source diet resembles a traditional media plan. According to 5W Research's 2026 analysis, Wikipedia accounts for 13.15 percent of ChatGPT citations in the U.S. and Reddit for 11.97 percent, which means two reference and community sites together supply over a quarter of everything ChatGPT cites. Meanwhile the Wall Street Journal, the New York Times and Bloomberg do not appear in the top 20 at all. Prestige with human editors and prestige with a retrieval pipeline are different currencies.

Larger datasets confirm the shape. Otterly's 2026 AI Citations Report, built on more than a million data points, and Semrush's three-month study of the most-cited domains both found the same pattern: Wikipedia and Reddit at the top, then a long, fragmented tail of niche sites, review platforms and individual pages that happen to answer specific questions well. That long tail is the opportunity. You will probably never be Wikipedia, but you can absolutely be the page in the tail that answers one buying question better than anyone else. We break the full source data down in what ChatGPT actually cites.

FindingNumberStudy
Wikipedia share of ChatGPT citations (U.S.)13.15%5W Research, 2026
Reddit share of ChatGPT citations (U.S.)11.97%5W Research, 2026
WSJ, NYT, Bloomberg in the top 20 cited domainsAbsent5W Research, 2026
ChatGPT-cited pages carrying structured data~71%SE Ranking
Domains cited by both ChatGPT and Perplexity~11%Cross-platform citation analyses (Profound)

The structured data figure is a correlation, not proven causation; see the schema section below for the conflicting evidence.

One more number belongs here because it reframes strategy: cross-platform analyses find only about 11 percent of domains are cited by both ChatGPT and Perplexity. Each engine has its own diet, so a single content strategy does not automatically win every AI surface. Optimizing for ChatGPT specifically means studying what ChatGPT specifically cites for your questions, which is why Reachroller stores the full answer text for every run instead of a score alone.

Why the same question returns different brands

Ask ChatGPT the identical buying question twice and you will often get two different brand lists. SparkToro measured this directly and found under a 1 percent chance that two identical ChatGPT runs return the same list of brands. This is a property of how the technology works: language models sample from probability distributions when they generate text, retrieval can surface a different page set run to run, and small differences compound into different shortlists. There is no bug to report and no stable ranking hiding underneath.

The measurement consequence is brutal for casual checking. A founder who asks ChatGPT once, sees their brand mentioned, and concludes the channel is handled has measured a coin flip. The reverse error is just as common: one bad run and a team panics about an answer that names them the majority of the time. The honest unit of measurement is frequency over repeated runs, tracked as a trend line. We unpack the volatility research fully in why AI gives a different answer every time.

This is also the fact that shapes Reachroller's methodology: questions are run repeatedly, a mention only counts when the brand name literally appears in the stored answer text, every number links back to the raw answer, and branded questions are excluded from the headline score because an answer to "is Acme any good" mentions Acme by construction. Whatever tool or manual process you use, insist on those properties, or the score is decoration.

What the research says moves recommendations

The strongest experimental evidence on influencing generative answers is still the Princeton and Georgia Tech GEO study, published at KDD 2024. The researchers tested nine optimization methods and found that visibility in generative engine responses can be boosted by up to roughly 40 percent. The best-performing techniques were adding quotations, adding statistics, and citing sources: the top methods improved about 22 percent on position-adjusted word count and about 37 percent on subjective impression versus baseline. In plain terms, pages that read like evidence get picked; pages that read like assertion get skipped.

Just as instructive is what failed. Keyword stuffing, a tactic with a long history in classic SEO, performed near the bottom for generative engines. The model is not matching strings, it is reading, and a page bloated with repeated phrases reads worse. The study also found effectiveness varies by domain, so the right move is testing techniques against your own category's questions rather than copying a generic checklist. This finding is the intellectual foundation of the whole discipline, and it is why every fix page Reachroller generates leads with quotable claims, sourced statistics and a structure built to be excerpted.

Then there is schema markup, where the honest summary is a conflict. SE Ranking found about 71 percent of ChatGPT-cited pages carry structured data. But Ahrefs' May 2026 experiment on 1,885 pages found that adding JSON-LD schema produced no measurable citation lift for ChatGPT on pages that were already cited. Both can be true: schema may matter for initial parsing and discovery while adding nothing once a page is established. Our position is pragmatic. Ship schema because it is nearly free, and never mistake it for the work of writing a better answer.

The revenue stakes, measured

If the mechanics feel abstract, the buyer data makes them concrete. In G2's 2026 research, 69 percent of B2B software buyers chose a different vendor than they originally expected because of AI chatbot output, and 33 percent bought from a brand they had never heard of before the AI named it. Read that second number again. A third of buyers are discovering their eventual vendor inside a chat window. The single most common use of AI in software research is comparing vendor strengths and weaknesses, at 41 percent, which is precisely the question format where ChatGPT names brands.

Now the uncomfortable statistic: G2 also found 51 percent of B2B tech brands have zero citations across ChatGPT, Perplexity and Gemini. Half the market is absent from the conversations deciding a growing share of deals. That absence is invisible in your analytics because the lost buyer never visited your site; the deal was shaped, and often settled, before any click. This is the case for measuring AI visibility at all, and we make it in full in what is AI visibility.

Forrester's numbers extend the pattern beyond software: 54 percent of buyers researched products with AI, and 47 percent built internal business cases with it before ever contacting a vendor. The recommendation you are trying to influence is not a top-of-funnel curiosity. It is feeding the document your champion shows their CFO.

How to influence ChatGPT's picks

Everything above compresses into a working sequence. First, find the questions you lose: run the 20 to 30 unbranded buying questions in your category through ChatGPT repeatedly and record which answers name you, which name rivals, and which sources get cited. Second, publish citable answers: for each lost question, create a page that answers it directly in the first paragraph, backed with the quotations, statistics and cited sources the GEO research validated. Third, get indexed: submit through Google Search Console and Bing Webmaster Tools, since retrieval reads from search indexes. Fourth, work the third-party layer, because the citation data shows independent sources carry most of the weight. Fifth, recheck after one to two weeks and measure the flip.

You can run that loop by hand, and our step-by-step playbook for getting mentioned by ChatGPT shows exactly how. The honest accounting is that the manual version costs hours per week: repeated runs, answer storage, mention counting, page drafting and recheck scheduling. Reachroller automates that loop end to end. It runs your questions through the official ChatGPT API on a schedule, scores mentions against stored answer text, generates a publish-ready fix page for every question you lose, complete with URL slug, title tag, meta description, schema markup and indexing steps, then rechecks the answer and shows whether it flipped.

Pricing is deliberately founder-sized: Starter is $29 per month for 400 credits and 25 tracked questions, where one credit is one AI answer and a generated fix costs ten. The trial is three days with 50 credits, every feature unlocked and no card required, which is enough to run a real first report on your own brand and generate one fix. The mechanics in this article decide which brands ChatGPT recommends. The only question left is whether you will be measuring them or guessing.

Frequently asked questions

Does ChatGPT recommend brands based on advertising?+

No. There is no paid placement inside ChatGPT's organic answers. Recommendations come from two organic mechanisms: patterns in the model's training data and pages retrieved through live web search. That is why the practical levers are citable content, third-party mentions and indexing rather than a media budget.

Can I pay OpenAI to have my brand mentioned?+

You cannot buy a mention in the answer itself. The only route is earning one: publishing content ChatGPT's retrieval can cite, appearing on the third-party sources it already trusts, and being present in the training data over time. Tools like Reachroller shorten the loop by finding the questions you lose and producing the pages that address them.

How long does it take to change what ChatGPT says about my brand?+

For answers grounded in live web search, one to two weeks is realistic: the page must be published, indexed by Google and Bing, then picked up by retrieval. For knowledge baked into the model's weights, changes arrive on OpenAI's retraining schedule, which you cannot influence or predict. This is why retrieval-driven questions are the ones worth working first.

Why does ChatGPT recommend my competitors and not me?+

Usually one of three reasons: your competitors appear on more of the third-party pages ChatGPT retrieves and cites, your own pages do not answer the buying question in a quotable way, or your brand simply has thinner coverage in the training data. Running the actual questions and reading the cited sources, which Reachroller stores for every answer, shows which reason applies.

Does schema markup make ChatGPT recommend a brand?+

The evidence is mixed. SE Ranking found roughly 71 percent of pages cited by ChatGPT include structured data, but that is correlation. An Ahrefs study of 1,885 pages in May 2026 found no measurable citation lift from adding schema to pages that were already cited. Ship schema because it is cheap and may help parsing, and put your real effort into answer quality and third-party mentions.

Is ChatGPT the only AI engine worth optimizing for?+

It is the biggest single target: ChatGPT commands about 92 percent of trackable LLM referral traffic and 72 percent of B2B software buyers use it during vendor evaluation, per G2. But cross-platform analyses find only about 11 percent of domains are cited by both ChatGPT and Perplexity, so winning one engine does not automatically win the others.

How does Reachroller measure ChatGPT recommendations?+

Through OpenAI's official API only, no scraping. It runs the unbranded buying questions in your category, counts a mention only when your brand name literally appears in the stored answer text, links every score to the raw answer, and excludes branded questions from the headline number because those mention you by construction.

Sources referenced

  • 5W Research, ChatGPT citation share analysis, 2026
  • Otterly.AI, The AI Citations Report 2026 (1M+ data points)
  • Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
  • Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
  • G2, B2B buyer AI research, 2026
  • SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
  • SE Ranking, ChatGPT referral traffic study, May 2026 (101,574 websites)
  • Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
  • Botify, analysis of OpenAI crawl growth, 2026; OpenAI developer docs on OAI-SearchBot

Find out which brands ChatGPT recommends instead of you.

Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.

Check my brand free