Research
The Princeton GEO paper, explained for marketers
Updated July 23, 2026
The GEO paper is the study that named generative engine optimization. Written by researchers at Princeton and Georgia Tech and published at KDD 2024 under arXiv:2311.09735, it tested how editing a web page changes that page's visibility inside AI-generated answers. The headline result: visibility in generative engine responses can be boosted by up to about 40 percent, and the best-performing techniques were adding quotations, adding statistics and citing sources, which improved roughly 22 percent on Position-Adjusted Word Count and 37 percent on Subjective Impression versus baseline. Keyword stuffing, the classic SEO reflex, landed near the bottom. Reachroller's generated fix pages are built around those findings, so the research is worth understanding in the original.
Why one academic paper matters this much
Most marketing disciplines accumulate slowly, out of practitioner folklore and vendor claims, and get their evidence base later if ever. Generative engine optimization ran in the opposite order. Before the agencies, the conference tracks and the job titles, there was a controlled study: researchers at Princeton and Georgia Tech asked whether editing a web page changes how visible that page is inside AI-generated answers, built a benchmark to test it, and published the results at KDD 2024. The paper is titled GEO: Generative Engine Optimization, it is archived as arXiv:2311.09735, and it gave the field both its name and its first falsifiable claims.
That founding matters practically, because it means GEO advice can be sorted into two piles: tactics with experimental support, and tactics someone made up. An industry now projected by Intel Market Research to reach about $1.48 billion in services revenue in 2026 sells plenty of both. Knowing what the original evidence actually says is the cheapest defense against buying the second pile.
This explainer covers what the researchers measured, which techniques won and lost, the caveats the paper itself flags, and what has and has not held up in the two years since. If you want the broader discipline that grew around the paper, start with our primer on generative engine optimization and come back for the source material.
The question the researchers asked
A generative engine answers a question by retrieving relevant pages, synthesizing them into a written response, and citing some of them as sources. From a website owner's perspective, this creates a new kind of competition: your page is no longer competing for a ranked slot on a results page, it is competing for space inside a composed answer. The paper's question follows directly: if a source page is rewritten in different ways, which rewrites earn it more presence in the generated answer?
To test this systematically, the authors built GEO-bench, a benchmark of queries spanning many domains, and ran optimization methods against it, comparing each optimized page's visibility in generated answers to its unoptimized baseline. The methods themselves were deliberately practical edits a content team could apply: adding quotations, adding statistics, citing sources, adjusting fluency and style, and including the old SEO standby, keyword stuffing, as a control from the previous era of search optimization.
This design is what separates the paper from opinion. Each technique gets the same queries, the same engines and the same scoring, so when one edit outperforms another, the difference is attributable to the edit. Two years of vendor decks have quoted the results; far fewer explain the measurement, which is where the interesting detail lives.
How the paper measures visibility
You cannot optimize what you have not defined, so the paper's first contribution is a pair of visibility metrics for generated answers. The first, Position-Adjusted Word Count, is the objective one: it measures how much of the answer's text draws on your page, weighted by position, so material that appears early in the answer counts for more than material buried at the end. It captures the intuition that being the answer's opening sentence is worth more than being its footnote.
The second metric, Subjective Impression, evaluates how the source comes across in the answer: how prominent, useful and favorably presented it is to a reader. The two metrics together anticipate a distinction that now runs through the whole industry, between being quoted at length and being framed well, roughly parallel to the split between citations and mentions that we unpack in our guide to the two currencies of AI visibility.
Note what this measurement philosophy implies: visibility is a property of stored answer text, measured against a baseline, over many queries. It is the same philosophy Reachroller adopted for brand tracking, where a mention only counts when the brand name literally appears in the stored answer and every score links back to the raw text. The paper set the standard that honest AI visibility measurement is corpus-level and auditable, and the tools worth using inherited it.
What won: quotations, statistics, cited sources
The best-performing techniques across the benchmark were adding quotations, adding statistics and citing sources. The strongest methods improved visibility by roughly 22 percent on Position-Adjusted Word Count and about 37 percent on Subjective Impression versus baseline, and the paper reports that visibility in generative engine responses can be boosted by up to about 40 percent. For a discipline where practitioners argue endlessly about tactics, having three editing moves with measured double-digit effects is remarkable.
Notice what the three winners share: they all increase evidence density. A quotation is a specific person saying a specific thing. A statistic is a measured quantity. A cited source is a verifiable trail. Each gives a synthesis engine something it structurally wants: concrete, attributable material to build an answer from. A language model composing a response needs claims it can compress and attribute, and a page full of vague assertions offers nothing to grab. The paper effectively measured how much engines prefer pages that read like evidence over pages that read like advertising.
This is also the finding that aged best. The citation studies that followed fit it neatly: 5W Research's 2026 analysis found Wikipedia, the web's densest evidence format, is ChatGPT's single largest source at 13.15 percent of U.S. citations, while prestige newspapers miss the top 20. The engines reward verifiable density, exactly as the paper predicted. How to build a page around that principle, section by section, is the subject of our guide to writing content AI engines cite.
What flopped: keyword stuffing and the SEO reflex
The paper's most quotable negative result: keyword stuffing performed near the bottom of the tested techniques for generative engines. The tactic that defined a generation of search optimization, repeating target phrases until the page chants them, does approximately nothing for AI answers. The mechanism is intuitive once stated. Classic ranking systems matched query strings against page strings, so string repetition could move relevance signals. A language model synthesizes meaning; it already knows your page is about project management software from reading it once, and the ninth repetition adds no quotable substance.
The strategic warning generalizes beyond one tactic: SEO instinct transfers incompletely to generative engines, and sometimes points backward. Habits built around keyword density, exact-match phrasing and page-one thinking can consume the exact effort that should go into quotations, statistics and sources. Which instincts carry over and which invert is the subject of our comparison of GEO and SEO.
To keep the caveat honest: indexing still gates everything, because generative engines retrieve from search indexes, and Google's AI features cite from Google's organic index. The paper does not license abandoning search fundamentals. It licenses abandoning the decorative parts of SEO while keeping the structural parts: crawlability, indexability and pages that genuinely answer questions.
The findings, in one table
| Technique | Result | Why it makes sense |
|---|---|---|
| Adding quotations | Among the best performers | Direct, attributable statements give engines liftable text |
| Adding statistics | Among the best performers | Concrete numbers make a page read as evidence |
| Citing sources | Among the best performers | Named references signal verifiability |
| Best methods combined effect | ~22% on Position-Adjusted Word Count, ~37% on Subjective Impression | Versus unoptimized baseline; up to ~40% overall boost |
| Keyword stuffing | Near the bottom | The classic SEO tactic does not transfer to generative engines |
Results as reported in arXiv:2311.09735 against the paper's GEO-bench benchmark; effects varied by domain.
The caveat the paper insists on: domain matters
The paper's most under-quoted finding is that technique efficacy varies by domain. The edits that lift a page answering historical questions differ from the edits that lift a product comparison or a technical explainer. The authors are explicit that domain-specific optimization matters, and GEO-bench itself spans many query categories precisely to surface this variation. In other words, the paper refutes its own worst summary: there is no single trick that maximizes visibility everywhere.
Later industry data extends the same lesson across engines. Cross-platform citation analyses find only about 11 percent of domains are cited by both ChatGPT and Perplexity, so what wins one engine mostly does not win another. Between domain variation and engine variation, blanket GEO advice is structurally suspect. The honest unit of optimization is one question on one engine, measured before and after.
That granularity is exactly how Reachroller structures its fix loop: each generated fix targets one lost buyer question, is written with the evidence-density findings applied to your category, ships with slug, title, meta description, schema markup and indexing steps, and gets a recheck to confirm whether that specific answer flipped. The paper's caveat, turned into product design.
What holds in 2026, and what to hold loosely
Two years is a long time in this field, so it is fair to ask what survives. The direction survives everything: engines still favor specific, verifiable, evidence-dense content, and every major citation study since fits that frame. The buyer context grew stronger, which raises the stakes on the findings: Forrester's 2026 survey found 94 percent of buyers used AI in their most recent purchase, and G2 found 69 percent of software buyers changed their expected vendor because of chatbot output. The answers the paper taught us to optimize now decide real purchases, as we detail in our unpacking of the B2B buyer data.
Hold the exact percentages loosely. The engines of 2024 are several product generations behind the engines of 2026, so 22, 37 and 40 are historical measurements of a moving system rather than constants. Hold single-run measurements even more loosely: SparkToro found under a 1 percent chance that two identical ChatGPT runs return the same brand list, so any before-and-after test needs repeated runs to mean anything. And note the questions the paper never addressed, where later evidence is genuinely mixed: structured data, for instance, correlates with citation in SE Ranking's data, at about 71 percent of ChatGPT-cited pages, while an Ahrefs experiment on 1,885 already-cited pages found no lift from adding schema. The paper is a foundation; it is nobody's complete playbook.
The fair summary after two years: the GEO paper made three claims a marketer can act on, evidence density lifts visibility, keyword stuffing does not, and effects depend on domain. All three have aged well. The industry built on top of them has produced plenty of noise, but the signal underneath is still the paper's.
Applying the paper on Monday morning
Here is the paper reduced to a working procedure. List the buyer questions that matter to your revenue, and find out which AI answers currently omit you, using repeated runs rather than one check. For each lost question, take the page that should win it, or create one, and rewrite it as evidence: an answer-first opening, real statistics with named sources, quotations from people with standing, and citations the engine can follow. Skip the keyword ritual entirely. Get the page indexed in Google and Bing, since retrieval reads the indexes, and then remeasure the same question over the following weeks.
You can run that loop by hand, and for a handful of questions you should at least once, because reading raw engine answers about your own category is irreplaceable education. At scale it becomes a job. Reachroller runs it as software: tracked questions across ChatGPT, Claude, Gemini, Perplexity and Grok through official APIs, scoring that only counts literal mentions in stored answers with branded questions excluded, a generated fix page per lost question built on the techniques the paper validated, and a recheck that tells you whether the answer moved. Starter is $29 per month, and the trial runs a full first report on your own domain.
Read the paper itself if you have an afternoon; it is unusually accessible for an academic venue, and arXiv:2311.09735 is free. Two years on, it remains the highest ratio of evidence to hype available in this field, and the tactics it validated are still the ones that move answers.
Frequently asked questions
What is the GEO paper?+
GEO: Generative Engine Optimization is the first academic study of how web content can be optimized for visibility inside AI-generated answers. It was written by researchers at Princeton and Georgia Tech, published at KDD 2024, and is available as arXiv:2311.09735. It coined the term GEO and introduced GEO-bench, a benchmark for evaluating optimization methods.
What did the GEO paper actually find?+
That editing a page changes how visible it is in generative engine responses, by up to about 40 percent. The best-performing techniques were adding quotations, adding statistics and citing sources, which improved roughly 22 percent on Position-Adjusted Word Count and about 37 percent on Subjective Impression versus baseline. Keyword stuffing performed near the bottom.
What is Position-Adjusted Word Count?+
One of the paper's two visibility metrics. It measures how much of a generative answer's text draws on your page, weighted by how early in the answer that material appears. Earlier and longer presence scores higher. The companion metric, Subjective Impression, evaluates how prominently and favorably the source comes across in the answer.
Does keyword stuffing work for AI search?+
No. In the GEO paper's experiments, keyword stuffing ranked near the bottom of tested techniques for generative engines. Language models synthesize meaning rather than match strings, so repeating target phrases adds nothing an engine wants to quote. Evidence density, quotations, statistics and named sources, is what moved visibility.
Do the GEO paper's findings still hold in 2026?+
Directionally, yes. The engines have evolved since the experiments, so treat exact percentages as historical. But the citation research that followed points the same way: engines favor specific, evidence-dense, verifiable content, and 2026 buyer studies from Forrester and G2 confirm the answers being optimized now shape real purchase decisions. The paper also found effects vary by domain, which later cross-engine studies echo.
How do I apply the GEO paper to my own site?+
Rewrite the pages that target your most valuable buyer questions to lead with a direct answer backed by statistics, quotations and named sources, then get them indexed and measure whether AI answers change. Reachroller operationalizes exactly this: it finds the questions you lose, generates a fix page built on the evidence-density findings, and rechecks the answer after you publish.
Sources referenced
- Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- 5W Research, ChatGPT citation share analysis, 2026
- SE Ranking, structured data on AI-cited pages (ChatGPT and Google AI Mode)
- Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
- Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
- G2, B2B buyer AI research, 2026
- Intel Market Research, GEO services market outlook, 2026
Put the GEO research to work on your own answers.
Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.
Check my brand free