Engines
How AI assistants choose sources: the cross-engine picture
Updated July 20, 2026
How do AI assistants choose their sources? Differently, engine by engine, and the differences are measured. ChatGPT leans on Wikipedia and Reddit, which together account for over a quarter of its U.S. citations according to 5W Research, while major newspapers miss its top 20. Perplexity leans even harder on community content and cites about 8.2 sources per answer. Google's AI surfaces cite from Google's organic index. Only around 11 percent of cited domains overlap between ChatGPT and Perplexity, so no single page wins everywhere. What travels across all engines is the content itself: indexed pages with statistics, quotations and cited sources, which the Princeton GEO research found lift visibility by up to roughly 40 percent. Reachroller shows you the exact sources each engine cited for the questions you lose.
Five engines, five diets
When an AI assistant answers a buying question, it is doing source selection before it does anything else: deciding which handful of pages, out of everything it knows and everything it can retrieve, deserves to shape the answer. That selection step is where brands are made visible or invisible, and it is no longer a black box. Across 2025 and 2026 a series of large citation studies, including Otterly's AI Citations Report built on more than a million data points and Semrush's three-month study of the most cited domains, mapped what each engine actually cites at scale.
Two findings organize everything else. First, the big studies agree on the headliners: Wikipedia and Reddit dominate AI citations, with a long, fragmented tail of niche sites behind them. Second, the engines disagree with each other: cross-platform analyses, including work by Profound, find only about 11 percent of domains are cited by both ChatGPT and Perplexity. The same question, asked of two engines, is answered from two mostly non-overlapping reading lists.
That combination, shared favorites plus divergent tails, is why this article goes engine by engine. A brand that understands each engine's diet can place its effort where that specific engine actually looks. A brand that treats AI visibility as one undifferentiated channel is optimizing for an average reader that does not exist.
ChatGPT: reference and community over legacy media
The clearest citation map we have belongs to the biggest engine. 5W Research found that Wikipedia accounts for 13.15 percent and Reddit for 11.97 percent of ChatGPT's citations in the U.S., together over a quarter of everything it cites. The absences are as telling as the leaders: the Wall Street Journal, the New York Times and Bloomberg do not appear in the top 20. When ChatGPT reaches for evidence, it reaches for reference material and practitioner conversation before prestige journalism.
The infrastructure behind the selection is shifting too. ChatGPT's search launched on Bing's index, but OpenAI now operates its own crawler, OAI-SearchBot, and has roughly tripled its web crawl since August 2025 according to Botify's analysis. OpenAI is increasingly deciding for itself what the web contains rather than inheriting Bing's view of it, which means being crawlable by OAI-SearchBot and present in the sources ChatGPT already trusts are both live levers. We break the full citation data down in what ChatGPT actually cites.
For a brand, the ChatGPT playbook follows the map: an accurate presence on the reference sites it leans on, genuine participation in the communities it reads, and owned pages built to be citable when its search retrieves them. Given ChatGPT's scale, this map is where most brands should start, and the one whose levers pay back first.
Perplexity: citation-rich, community-heavy
Perplexity is the engine most transparent about its sources, because citations are its product. It averages about 8.2 sources per answer, roughly 3.4 times ChatGPT's habit, drawing on its own index reported at more than 50 billion pages. More citation slots per answer arithmetically widens the door: a page does not need to be the single best source on the web to be cited by Perplexity, it needs to be among the best eight for that question.
Its diet skews harder toward community content than any other engine. Reddit is Perplexity's single largest source, with estimates ranging from about 17 to 24 percent of citations, and one analysis put Reddit at 46.7 percent of Perplexity's top-10 citation share. The spread between those figures is itself informative: methodologies and query mixes differ, so quote ranges, never point estimates. Beyond Reddit, Perplexity skews toward LinkedIn, NIH and G2, a profile that rewards professional presence and review-site standing. The engine-specific playbook is in how Perplexity picks its sources.
Perplexity matters beyond its query volume because of when buyers use it. G2's 2026 research found 44 percent of B2B software buyers use Perplexity during shortlisting, the stage where vendor lists get cut. An engine that cites eight sources at the exact moment a shortlist forms is, per unit of effort, the most winnable citation real estate in the field.
Google's surfaces: the organic index is the door
Gemini, AI Overviews and AI Mode select sources the most legible way of all: from Google's organic index. Whatever Google's crawling, indexing and quality systems admit is eligible; whatever they exclude does not exist for these surfaces. There is no separate AI submission channel and no workaround, which makes the first move on Google's surfaces identical to the oldest move in search: be indexed, be crawlable, be worth retrieving.
Within the eligible pool, the selection favors pages that answer precisely. On structured data, report the conflict honestly: SE Ranking found about 71 percent of pages cited by ChatGPT and 65 percent cited by Google AI Mode carry structured data, a correlation, while Ahrefs' May 2026 study of 1,885 pages found adding JSON-LD produced no measurable citation lift on already-cited pages and a statistically significant decline for AI Overviews. The defensible read is that schema helps machines parse you and still earns rich results, but content, an early complete answer with attributed numbers, is what earns the citation.
One practical note that surprises people: because the three Google surfaces share an index, source selection work done for one compounds across all three. A page that Gemini grounds on is eligible for the AI Overview and for AI Mode without additional effort, which makes Google the most efficiency-friendly of the engine families even when it is not the first priority.
Claude and Grok: where the studies run out
The citation studies thin out fast beyond the big three, so intellectual honesty means switching from measured shares to mechanics. Claude answers from training data plus a web search it runs on demand, and its selection habit is temperamental rather than statistical: it hedges, names fewer brands per answer, and favors sources it can quote with confidence. That raises the bar for inclusion and raises the value of honest, trade-off-aware pages. The full picture is in how Claude handles brand questions.
Grok adds a source no other engine touches: live posts on X, alongside web search and training data. Its selection can include what practitioners said about your category this week, which makes it the fastest-moving surface in AI visibility and the most volatile. Brands in X-active categories get a lane there that the static web never offered; brands elsewhere are covered by their standard web work. We map the mechanics in Grok and brand visibility.
The cross-engine picture, in one table
| Engine | Leans on | Citation habit | First move for brands |
|---|---|---|---|
| ChatGPT | Wikipedia (13.15%), Reddit (11.97%), reference and community sites | Fewer citations per answer; own crawler growing fast | Citable pages plus presence on reference and community sites |
| Perplexity | Reddit heaviest of all, plus LinkedIn, NIH and G2 | ~8.2 sources per answer, roughly 3.4x ChatGPT | Earn community mentions; publish precise, quotable answers |
| Google (Gemini, AI Overviews, AI Mode) | Google's own organic index | Synthesis across retrieved pages; classic quality signals apply | Be indexed and rankable; answer the question in paragraph one |
| Claude | Web search on demand; no published citation-share studies | Hedged answers, fewer named brands per response | Honest, trade-off-aware pages it can quote with confidence |
| Grok | Live X posts plus web search; least studied engine | Fast-moving answers shaped by current discourse | Genuine X participation plus the standard citable pages |
Citation shares from 5W Research and cross-platform analyses, 2026; Claude and Grok characterized qualitatively where measured shares do not exist.
Reddit's rise and the fragmented long tail
One trend line runs through every engine's diet: community content is gaining. Reddit's AI citation share grew about 73 percent in commercial categories across 2025 and 2026, and it now sits in the top two sources for both ChatGPT and Perplexity. The logic is easy to reconstruct: when a buyer asks which product to choose, threads where practitioners compare real experiences are closer to the answer than anything a brand publishes about itself. Engines select for usefulness, and communities manufacture usefulness at scale. What that means for brand participation, and why astroturfing backfires, is covered in our analysis of Reddit's AI citation rise.
Below the headliners, the citation studies describe a long, fragmented tail: niche review sites, category blogs, comparison pages and reference resources that each hold a sliver of citation share but collectively decide most answers. This fragmentation is good news for smaller brands. You do not need the Wall Street Journal to notice you, since the engines barely cite it anyway. You need to be present, accurately and favorably, on the two dozen specific pages each engine already trusts for your category's questions. Those pages are findable, and pitching them is ordinary work once you know which ones they are.
What earns a pick on every engine
Under the divergent diets sits a shared selection instinct, and it has been measured. The Princeton-led GEO study, published at KDD 2024, tested nine optimization methods across generative engines and found that adding quotations, statistics and cited sources performed best, lifting visibility by up to roughly 40 percent, with the strongest methods improving about 22 percent on position-adjusted word count and about 37 percent on subjective impression versus baseline. Keyword stuffing, the reflex imported from old SEO, performed near the bottom. Engines select evidence, and pages that look like evidence get selected.
The study also found efficacy varies by domain, which is a polite way of saying: test on your own questions. The transferable core is structural. A citable page states the complete answer in its first paragraph, attributes every number to a named source, uses question-shaped headings a retrieval system can match, stays honest about trade-offs, and is indexed everywhere engines look. Add the off-page half, being named by sources the engine already trusts, and you have the whole cross-engine playbook in two sentences. Everything else is engine-specific tuning.
Watching source selection happen on your own questions
Everything above describes averages across millions of answers, and your brand does not live in the average. The question that pays your bills is which sources each engine cites for your buyers' twenty questions, and the only way to know is to ask, repeatedly, and read the citations. This is the workflow Reachroller automates: it runs your tracked questions through official engine APIs, stores every raw answer with the sources it cited, and shows you the actual reading list behind every answer you win or lose. The method is documented on our methodology page.
Knowing the reading list changes what you do next. When the engine cites a comparison page that omits you, the move is outreach to that page. When it cites nothing recent and precise, the move is publishing the page that fills the gap, and Reachroller generates it for every lost question, publish-ready with slug, title tag, meta description, schema markup and indexing steps, then rechecks the answer to show whether it flipped. That loop, from observed source selection to published change to verified result, is the entire discipline of AI visibility compressed into a workflow, and it starts at $29 per month with a three-day, no-card trial.
How question type changes source selection
Engine diets are averages, and averages hide a second pattern: the same engine reaches for different shelves depending on what kind of question it is answering. Comparative buying questions, which tool is best, X versus Y for a given team, pull hardest on community threads, review platforms and third-party comparison pages, because those are where trade-off judgments live. Perplexity's skew toward G2 and Reddit is the measured version of this instinct. Factual and specification questions, what something costs, whether a product integrates with another, pull toward vendor documentation and reference pages, where a precise answer can be lifted verbatim.
High-stakes domains shift the selection again. Perplexity's measured skew toward NIH is the clearest example: for health-adjacent questions, engines reach for institutional authority over community opinion, and the same caution shows up around financial and legal questions. And anything time-sensitive pulls toward fresh pages and, on Grok, live discourse. For a brand, this maps question types to content types with unusual precision: your documentation must carry the facts an engine lifts for spec questions, third parties must carry the judgments it lifts for comparisons, and no single page can do both jobs.
This is why a serious visibility effort audits its question list by type before it writes anything. Sort your twenty buyer questions into factual, comparative and situational piles, and you will usually find the losses cluster: strong documentation but no third-party comparisons, or glowing reviews but documentation too vague to quote. The cluster tells you what to build next, and it is a diagnosis a source-blind visibility score can never give you, which is exactly why Reachroller stores the cited sources beside every answer instead of reducing the run to a number.
Frequently asked questions
What sources does ChatGPT cite most?+
Wikipedia and Reddit, by a wide margin. 5W Research found Wikipedia at 13.15 percent and Reddit at 11.97 percent of ChatGPT's U.S. citations, together over a quarter of the total, while the Wall Street Journal, New York Times and Bloomberg do not appear in the top 20. Reference and community content beats legacy media in ChatGPT's diet.
Why does Perplexity cite so many more sources than ChatGPT?+
It is built as an answer engine with citations as the product. Perplexity averages about 8.2 sources per answer, roughly 3.4 times ChatGPT's habit, drawing on its own index reported at more than 50 billion pages. More citation slots per answer means more openings for a well-made page, which is why B2B brands often earn their first AI citation on Perplexity.
Do all AI engines use the same sources?+
No, and the divergence is the key strategic fact. Cross-platform citation analyses find only about 11 percent of domains are cited by both ChatGPT and Perplexity. Each engine reads the web through different infrastructure and different habits, so a brand needs to check its standing engine by engine rather than assuming one result generalizes.
Does structured data help a page get cited?+
The evidence conflicts, and honesty requires saying so. SE Ranking found about 71 percent of pages cited by ChatGPT and 65 percent cited by Google AI Mode carry structured data, which is correlation. Ahrefs tested 1,885 pages in May 2026 and found adding JSON-LD produced no measurable lift on already-cited pages and a significant decline for AI Overviews. Ship schema for parsing and rich results, and put your citation hopes in content instead.
What kind of content do AI engines actually cite?+
The Princeton GEO study, published at KDD 2024, found adding quotations, statistics and cited sources were the best-performing techniques, lifting visibility in generative engine responses by up to roughly 40 percent, while keyword stuffing performed near the bottom. Pages that state a complete answer early, attribute numbers to named sources and stay honest about trade-offs match what every engine selects for.
Why is Reddit so important to AI engines?+
Community answers are specific, experience-based and abundant, which is what answer synthesis needs. Reddit is one of ChatGPT's top two cited domains per 5W Research, Perplexity's single largest source by most estimates, and its AI citation share grew about 73 percent in commercial categories across 2025 and 2026. Engines trust what practitioners say to each other more than what brands say about themselves.
How do I find out which sources AI engines cite for my buying questions?+
Ask the engines your buyers' questions and record the citations, repeatedly, because answers vary run to run. Reachroller automates this: it runs your tracked questions through official engine APIs, stores every answer with its cited sources, and shows exactly which pages the engine trusted, so you know which sites to pitch and what to publish. The trial is three days, 50 credits, no card.
Sources referenced
- 5W Research, ChatGPT citation share analysis, 2026
- Otterly.AI, The AI Citations Report 2026 (1M+ data points)
- Semrush, most-cited domains in AI, 3-month study, 2025-2026
- SE Ranking, structured data on AI-cited pages
- Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
- Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
- Botify, analysis of OpenAI crawl growth, 2026; OpenAI developer docs on OAI-SearchBot
- Profound and other cross-platform citation overlap analyses, 2026
- G2, B2B buyer AI research, 2026
See the exact sources AI engines trust for your buyers' questions.
Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.
Check my brand free