Research
What ChatGPT actually cites: the data behind the answers
Updated July 24, 2026
ChatGPT cites Wikipedia more than any other source, at 13.15 percent of U.S. citations, with Reddit close behind at 11.97 percent, according to 5W Research's 2026 analysis. Together those two sites account for over a quarter of everything ChatGPT cites, while the Wall Street Journal, the New York Times and Bloomberg miss the top 20 entirely. Beyond the head, citations fragment into a long tail of niche, specific pages, which is exactly where an individual brand can realistically earn a spot. Only about 11 percent of domains are cited by both ChatGPT and Perplexity, so winning one engine does not win the others. Reachroller tracks which sources each engine cites for your buyers' questions, so you know where to aim.
Why citations are the metric that matters
When ChatGPT answers a buying question with live search behind it, it does two things that matter to a brand. It names products in the answer text, and it cites the pages that informed the answer. The citations are the paper trail. They reveal which corners of the web the engine treats as trustworthy for which topics, and by extension, where a brand needs to exist to influence future answers. If your category's answers are consistently sourced from Reddit threads and two niche review sites, that list is your media plan.
Citations also matter because of where ChatGPT sits in the market. It commands roughly 92 percent of trackable LLM referral traffic, so its sourcing habits shape more brand discovery than every other assistant combined. In 2026 several research teams mapped those habits at scale: 5W Research analyzed ChatGPT citation shares in the U.S., Semrush ran a three-month study of the most-cited domains in AI answers, and Otterly published its 2026 AI Citations Report built on more than a million data points. Their findings converge on a picture that surprises most marketers, and this article walks through it study by study.
One framing note before the data. A citation and a mention are different events: a mention is your brand named in the answer text, a citation is your page linked as a source, and each has its own mechanics. This piece is about citations; the distinction itself gets a full treatment in our guide to mentions versus citations.
The headline finding: Wikipedia and Reddit run the table
5W Research's 2026 analysis of ChatGPT citations in the U.S. found Wikipedia at 13.15 percent and Reddit at 11.97 percent. Two sites, over a quarter of all citations. Nothing else comes close, and the shape of the curve after them is a steep drop into a long tail of specific, niche pages. For anyone who spent a career equating authority with brand-name publishers, the composition of that top two is the real news: an encyclopedia anyone can edit and a forum where anyone can post outrank every newsroom in the country.
The absences are as telling as the leaders. The Wall Street Journal, the New York Times and Bloomberg do not appear in ChatGPT's top 20 cited sources in the 5W analysis. The publications that dominate human media consumption, win the awards and set the news agenda are minor sources in the machine's diet. Whatever the engine is optimizing for when it selects sources, prestige is visibly not the primary variable.
And the pattern replicates. Semrush's three-month most-cited-domains study and Otterly's 2026 AI Citations Report both confirm the same two findings: Wikipedia and Reddit dominance at the head, and long-tail fragmentation beyond it. When three independent datasets built on different query panels agree on the shape of the distribution, you can treat the shape as real even while holding the exact percentages loosely.
Why the machine trusts what it trusts
The studies measure what gets cited, and none of them can fully explain why, so treat this section as informed reading of the evidence rather than settled fact. Several properties plausibly favor Wikipedia and Reddit. Both are freely crawlable at enormous scale, while many major news sites meter or wall their archives, which matters to an engine assembling sources through live retrieval. Both are structured around questions and topics rather than news cycles: a Wikipedia article is a permanent, consolidated reference on one subject, and a Reddit thread is often a direct answer to the exact question a user just asked, written by people with firsthand experience.
There is also a format argument. Answer engines need text they can lift, attribute and compress. Encyclopedic prose and threaded discussion both compress well: they contain direct claims, comparisons and conclusions rather than narrative scene-setting. A news feature that spends four paragraphs on atmosphere before its first fact gives a retrieval system little to grab. Whatever the causal mix, the practical direction is identical: the engine favors consolidated reference content and authentic community discussion, and a brand can position itself in both.
One measured data point supports the format story. SE Ranking found that about 71 percent of pages cited by ChatGPT include structured data, and about 65 percent for Google AI Mode. That is a correlation, and the study authors are careful to say causation is unproven. An Ahrefs experiment in May 2026, covering 1,885 pages, found that adding JSON-LD schema produced no measurable citation lift on pages that were already cited. Machine-legible structure travels with citability without demonstrably causing it, a nuance we examine in our evidence review on schema markup for AI search.
ChatGPT versus Perplexity: two different diets
It is tempting to assume the AI engines all cite roughly the same web. The data says otherwise. Perplexity leans much harder on community content than ChatGPT does: Reddit is its single largest source, with estimates ranging from about 17 to 24 percent of its citations, and one analysis put Reddit at 46.7 percent of Perplexity's top-10 share. Perplexity also skews toward LinkedIn, NIH and G2, a mix that reflects its positioning as a research tool. And it cites more generously: about 8.2 sources per answer on average, roughly 3.4 times ChatGPT's citation density, which means more open doors per answer for a brand trying to get in.
The most strategically important number in this entire literature is the overlap figure: only about 11 percent of domains are cited by both ChatGPT and Perplexity, according to cross-platform citation analyses. Nine out of ten domains that win one engine do not win the other. A single content strategy cannot cover every AI surface, and any tool or agency promising blanket AI visibility from one tactic is contradicted by the data. Engine choice comes first, and it should follow your buyers: we walk through that prioritization in our guide to which AI engines matter for your brand.
This divergence is also why Reachroller tracks engines separately rather than blending them into one score. A brand can be strong in ChatGPT answers and invisible in Perplexity's, and an averaged number would hide exactly the gap you need to see. Each tracked question stores the answer per engine, with the cited sources attached, so the differences stay visible.
The Perplexity skew has a concrete B2B implication worth pausing on. If your buyers shortlist there, your review-site profiles and professional network presence are not side projects; they are the raw material the engine quotes. A software brand with a thin G2 profile is thin in exactly the place Perplexity looks first, and no amount of on-domain publishing compensates for an absence on the surfaces a given engine trusts.
Reddit's rise is the trend to watch
Within the citation data, one line is moving faster than the rest. Reddit's AI citation share grew about 73 percent in commercial categories across 2025 and 2026. Commercial categories are the ones with money attached: product comparisons, best-of questions, buying advice. Precisely where brands compete hardest, the engines are reaching more often for community threads written by people with no marketing budget and no editor.
The mechanism is easy to sympathize with. When a buyer asks which tool is best for a specific job, a thread of practitioners comparing notes is often the most honest document on the internet about that question. The engines appear to have learned what human searchers learned years ago when they started appending the word reddit to Google queries. For brands, the implication is double-edged: authentic community presence now feeds AI answers directly, and inauthentic presence carries platform bans and reputational blowback. The participation approach that compounds instead of backfiring is the subject of our Reddit playbook for brands.
Wikipedia is the other side of the same coin: the single largest ChatGPT source, with strict notability rules and an editorial community hostile to promotion. Brands that qualify for coverage benefit enormously; brands that try to force it get burned publicly. Neither surface can be bought, which is arguably why the engines trust them.
The Google comparison: same physics, different index
ChatGPT's citation habits matter most because of its referral dominance, but Google's answer surfaces run on the same physics with one structural difference: Google's AI features cite pages from Google's own organic index. Being indexed and rankable remains a precondition there, which means the classic search fundamentals you already maintain still buy you standing. The scale is substantial: AI Overviews trigger on roughly 48 percent of tracked queries by early 2026 according to third-party tracking, though some panels report 13 to 25 percent depending on the query set, and Google's Gemini app has surpassed 750 million monthly users.
The click math on those surfaces makes citation status decisive. Organic click-through drops about 61 percent when an AI Overview is present, according to 2026 tracking studies, while brands cited inside the overview see roughly 35 percent higher click-through. On a results page where the answer absorbs most of the attention, being one of the answer's named sources is the difference between losing the click and inheriting it.
So a complete citation strategy runs on two tracks. For Google's surfaces, the path goes through pages that already rank, since the overview is assembled from the organic index. For ChatGPT and Perplexity, the path goes through the source diets described above: reference content, community presence and the specific long-tail pages each engine retrieves. The tracks overlap in craft, both reward direct and evidence-dense pages, but they are measured separately, and a brand can be strong on one and absent from the other without ever noticing.
The citation data, in one table
| Finding | Number | Study |
|---|---|---|
| Wikipedia share of ChatGPT citations (U.S.) | 13.15% | 5W Research, 2026 |
| Reddit share of ChatGPT citations (U.S.) | 11.97% | 5W Research, 2026 |
| WSJ, NYT, Bloomberg in ChatGPT's top 20 sources | Absent | 5W Research, 2026 |
| Reddit share of Perplexity citations | ~17 to 24% (one analysis: 46.7% of top-10 share) | Cross-platform citation analyses |
| Average sources cited per Perplexity answer | ~8.2 (roughly 3.4x ChatGPT) | Citation density analyses, 2026 |
| Domains cited by both ChatGPT and Perplexity | ~11% | Cross-platform citation analyses |
| Reddit citation growth in commercial categories | ~73% (2025-2026) | Citation tracking studies |
| Cited ChatGPT pages carrying structured data | ~71% | SE Ranking (correlation, not causation) |
Percentages vary by panel and time window; the pattern, Wikipedia and Reddit at the head with a fragmented long tail, is consistent across studies.
The long tail is where brands actually win
It would be easy to read the head of the distribution and despair: you are unlikely to become Wikipedia, and you cannot own Reddit. But over 70 percent of ChatGPT's citations, on the 5W numbers, go to everything else, and that everything else is fragmented across niche sites, specific comparison pages, documentation, review platforms and focused blogs. Fragmentation is opportunity. The engine is not looking for the biggest site on a topic; the evidence suggests it is looking for the page that most directly answers the specific question in front of it.
That reading aligns with the strongest experimental evidence in the field. The Princeton and Georgia Tech GEO paper, published at KDD 2024, found that adding quotations, statistics and cited sources boosted a page's visibility in generative engine responses by up to about 40 percent, while keyword stuffing landed near the bottom. Pages that read like evidence get picked; pages that read like advertising do not. The practical craft of building such pages, section by section, is covered in our guide to writing content AI engines cite.
Two operational notes complete the picture. First, ChatGPT can only cite what its retrieval reaches: OpenAI operates its own crawler, OAI-SearchBot, and has roughly tripled its web crawl since August 2025 per Botify's analysis, but indexability still gates everything, so a page invisible to search indexes is invisible to the engine. Second, citation behavior is probabilistic. SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list, so a source cited in one run may be absent in the next. Single checks mislead; repeated measurement is the only honest read.
Turning citation data into a plan
Here is the workflow the data supports. Start by finding out what ChatGPT cites for your questions specifically, because category-level statistics are a map and your buyers' questions are the territory. Run the 20 to 25 questions your buyers ask, record which brands each answer names and which sources it cites, and repeat the runs over days rather than trusting one pass. The cited-source list that emerges is your target list: some entries will be pages you can build better versions of on your own domain, and some will be third-party surfaces where you need to earn a mention.
Then work both lists. On your own domain, publish answer-shaped pages built the way the GEO research rewards: direct claims, statistics, quotations, named sources. Off your domain, pursue the specific review sites, communities and reference pages the engine already trusts for your category, because getting named on a page the engine already cites moves answers faster than almost anything else. Reachroller automates the measurement half of this loop: it runs your questions on schedule through official APIs, stores every answer with its citations, counts a mention only when your brand name literally appears in the stored text, and generates a publish-ready fix page for each question you lose, then rechecks whether the answer flipped. The target list stops being a quarterly research project and becomes a live dashboard.
The citation studies will keep shifting as the engines evolve their sourcing. The durable lesson underneath them does not: AI engines cite the web they can read and the sources their users would trust. Be present, be specific, be verifiable, and be there before your rival is.
Frequently asked questions
What sources does ChatGPT cite most?+
Wikipedia leads at 13.15 percent of U.S. citations and Reddit follows at 11.97 percent, per 5W Research's 2026 analysis. Together they account for over a quarter of ChatGPT citations. Beyond those two, citations spread across a long tail of niche and specific pages rather than concentrating in a handful of big publishers.
Does ChatGPT cite major news sites?+
Far less than their cultural weight suggests. The Wall Street Journal, the New York Times and Bloomberg do not appear in ChatGPT's top 20 cited sources in 5W Research's 2026 U.S. analysis. Reference content and community discussion outrank prestige journalism in the machine's diet.
How is Perplexity's citation behavior different from ChatGPT's?+
Perplexity cites more sources per answer, about 8.2 on average, roughly 3.4 times ChatGPT, and it leans much harder on community content. Reddit is its single largest source, with estimates ranging from 17 to 24 percent of citations, and it also skews toward LinkedIn, NIH and G2. Only about 11 percent of domains are cited by both engines.
Does schema markup make ChatGPT cite a page?+
The evidence is correlational and contested. SE Ranking found about 71 percent of pages cited by ChatGPT include structured data, but an Ahrefs study of 1,885 already-cited pages in May 2026 found no measurable citation lift from adding JSON-LD schema. Schema may help initial parsing and discovery; it is not a citation switch.
How can my brand become a ChatGPT citation source?+
Two routes, and you likely need both. First, publish answer-shaped pages on your own domain for the specific questions buyers ask, since the citation long tail rewards specificity. Second, earn presence on the third-party surfaces ChatGPT already trusts, particularly Wikipedia where you qualify and Reddit through genuine participation. Reachroller shows which sources ChatGPT cited for each question you track, so you can target the surfaces that already win.
Does Google cite the same sources as ChatGPT?+
No, and the difference is structural. Google's AI features cite pages from Google's organic index, so ranking remains a precondition for being cited there, while ChatGPT retrieves through its own search stack. SE Ranking found about 65 percent of pages cited by Google AI Mode carry structured data versus about 71 percent for ChatGPT, a similar correlation on two different indexes. Measure each surface separately.
Do the citation studies agree with each other?+
On the headline pattern, yes. 5W Research, Semrush's three-month most-cited-domains study and Otterly's 2026 AI Citations Report, built on over a million data points, all confirm Wikipedia and Reddit dominance with long-tail fragmentation beyond them. They differ on exact percentages because they sample different query sets at different times, which is normal for panel research.
Sources referenced
- 5W Research, ChatGPT citation share analysis, 2026
- Otterly.AI, The AI Citations Report 2026 (1M+ data points)
- Semrush, most-cited domains in AI, 3-month study, 2025-2026
- SE Ranking, structured data on AI-cited pages (ChatGPT and Google AI Mode)
- Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
- Cross-platform citation analyses of ChatGPT and Perplexity domain overlap (Profound and others)
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- Botify, analysis of OpenAI crawl growth, 2026; OpenAI developer docs on OAI-SearchBot
See which sources ChatGPT cites for your buyers' questions.
Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.
Check my brand free