Playbooks
The anatomy of a page AI engines cite: a template
Updated August 2, 2026
Pages that AI engines cite share a repeatable skeleton of six blocks, in order: a standalone answer block of roughly 120 words directly under the H1; evidence paragraphs built on statistics with named sources and quotations, the levers the Princeton GEO study found lift generative visibility by up to 40 percent; a comparison or data table; a question-shaped H2 and H3 hierarchy; an FAQ section of 5 to 7 real buyer questions; and clean metadata with schema markup treated as hygiene rather than a lever, since Ahrefs measured no causal citation lift from adding schema alone. This article is the structural template; Reachroller generates pages in exactly this shape as 10-credit fixes.
Why structure decides citations
When a generative engine answers a buying question, it retrieves candidate pages, splits them into passages, and composes from the passages that read like answers. That pipeline has a consequence most content teams have not absorbed: the engine never grades your page as a whole. It grades chunks. A page can be authoritative, comprehensive and well ranked, and still contribute nothing to the composed answer because no single passage of it stands alone as something quotable. Structure, in other words, is not presentation. It is the difference between being retrieved and being used.
The evidence base for what works is unusually clean for a young discipline. The Princeton GEO study (KDD 2024) tested nine optimization methods across thousands of queries and found that adding statistics, quotations and cited sources lifted visibility in generative answers by up to 40 percent, while keyword stuffing fell below the do-nothing baseline. This template turns those findings, plus what 2026 citation data shows about tables, FAQs and schema, into an ordered skeleton you can apply to any page. The template also survives contact with the engines' differences: citation diets vary sharply between platforms, with only about 11 percent of cited domains overlapping between ChatGPT and Perplexity in cross-platform analyses, but the structural preferences, standalone answers, attributed evidence, extractable structure, hold across all of them, which is what makes a single template worth standardizing on. One scope note before the blocks: this is the structural template. The sentence-level craft, how to phrase claims, what to cut, why honesty outperforms, lives in our companion guide, how to write content AI engines actually cite. Use that one for the words and this one for the bones.
The template at a glance
| Block | Purpose | Spec |
|---|---|---|
| 1. Answer-first block | Gives retrieval a quotable passage where it looks first | 80 to 120 words under the H1, answers the title fully, stands alone out of context |
| 2. Evidence paragraphs | Supplies the statistics, quotes and sources engines extract | One claim per paragraph, each with a number and a named source |
| 3. Comparison or data table | Structured facts engines lift into their own comparisons | 3 to 8 columns, plain values, a caption naming the evidence |
| 4. Question-shaped headings | Matches passage retrieval to the questions buyers ask | H2s and H3s phrased as questions or direct claims, one topic each |
| 5. FAQ section | Catches long-tail phrasings of the same intent | 5 to 7 real questions, answers of 40 to 90 words, complete on their own |
| 6. Metadata and schema | Parsing and indexing hygiene | Unique title and description, canonical URL, Article and FAQPage JSON-LD |
Evidence: Princeton GEO study (KDD 2024) for blocks 1 and 2; Ahrefs May 2026 schema study for the block 6 caveat.
Block 1: the answer-first block
Directly under the H1, before any warm-up, write 80 to 120 words that fully answer the question the page exists for. The test is brutal and useful: could an engine lift this block verbatim, present it as the answer, and satisfy the asker? That means the block must carry its own context (name the entities, do not say "this tool" or "the approach above"), include the page's single most important number, and commit to a conclusion rather than promising one below. If the page is this article, the block states the six blocks. If the page is a pricing comparison, the block names the prices.
Two craft notes. First, write the block last, after the page has taught you what it concludes, then place it first. Second, resist the marketing urge to tease. A block that withholds the answer to force a scroll reads, to a retrieval system, like a page with no answer, and the citation goes to a competitor who committed. Every article on this blog, including this one, opens with exactly this block, so the pattern is visible one scroll up.
Block 2: evidence paragraphs, the measured lever
The body of the page is a sequence of evidence paragraphs, and the spec comes straight from the Princeton findings: statistics, quotations and cited sources are the three interventions with measured lift, up to 40 percent, so every important claim gets a number and a named source. "Adoption is growing fast" is filler; "Forrester's 2026 survey of 18,000 buyers found 55 percent compared vendors inside AI tools" is extractable. One claim per paragraph keeps passages self-contained, which matters because passage boundaries are where retrieval cuts.
Attribution style matters more than most writers expect. Engines prefer repeating claims that arrive pre-attributed, because the attribution transfers liability: "according to Ahrefs" is a sentence a cautious model can repeat safely. Name the organization, the year, and where possible the sample size. If a claim has no source you can name, the honest move is to cut it, and honesty is itself strategy here: invented numbers get repeated, then checked, then remembered. A closing sources list, like the one at the bottom of this page, concentrates the attributions where engines and skeptical readers both look.
Block 3: the comparison table
At least one table per page, because engines build comparisons and a table is comparison data in its native format. Forrester found 55 percent of buyers compared vendors inside AI tools during their most recent purchase; when an engine composes that comparison, clean rows with plain values are what it can lift without error. Keep cells short and factual (prices, counts, yes and no, one-line descriptions), keep the first column as the entity being compared, and add a one-line caption under the table naming the evidence behind it, which doubles as the attribution an engine can carry.
Tables also discipline the writer. A claim that cannot survive being a cell, stripped of adjectives, was probably an assertion rather than a fact. For the specific craft of pages built around head-to-head comparisons, where the table is the whole argument, see comparison pages and AI answers.
Block 4: question-shaped headings
Headings are the page's index for retrieval. Buyers ask engines full questions, and retrieval matches those questions against passages, so H2s and H3s phrased the way buyers phrase intent give the matcher an easy job. "How much does GEO cost per month?" outperforms "Pricing considerations" because it is the query. The discipline: one topic per heading, a hierarchy that never skips levels, and a section under each heading that answers its own heading completely, since any section may be read alone.
This is also where the keyword-stuffing instinct goes to die. Repeating the target phrase across every heading made sense to an older matching algorithm; to a generative engine it reads as noise, and the Princeton study measured stuffed content below the do-nothing baseline. Write headings for the question, once, clearly.
Block 5: the FAQ section
Five to seven questions, phrased the way real buyers ask them, each with an answer of 40 to 90 words that stands entirely on its own. The FAQ block earns its place by catching the long tail: the same intent arrives as "is GEO worth it", "does GEO actually work", and "should I pay for GEO", and a good FAQ covers phrasings the main sections did not. Source the questions from reality rather than imagination: sales calls, support tickets, community threads, and the queries your tracking shows you losing.
Resist stuffing the FAQ with restatements of sections above; engines deduplicate, and a FAQ that repeats the page adds nothing. Each answer should contain at least one concrete fact, and the best single question is the objection your sales calls hear most. The deeper practice of building pages around question sets is covered in FAQ pages for AI.
Block 6: metadata and schema, honestly weighted
The last block is hygiene, and it should be weighted as hygiene. Unique title tag that states the question and the answer's shape; meta description of 140 to 160 characters that a human would click; canonical URL; and Article plus FAQPage JSON-LD mirroring the visible content. Then stop. Ahrefs' May 2026 study tracked 1,885 pages that added schema against 4,000 controls and found citation changes within statistical noise: AI Mode +2.4 percent, ChatGPT +2.2 percent, AI Overviews -4.6 percent. Schema helps machines parse what is there; it does not make weak content citable, and weeks budgeted for markup are weeks taken from evidence.
The genuinely load-bearing hygiene is upstream: the page must be crawlable and indexed, because engines retrieve from search indexes and OpenAI's OAI-SearchBot, whose crawl Botify reports roughly tripled since August 2025. Check robots.txt, submit the URL, confirm indexing before judging results. The full markup picture, including where schema does still earn its keep, is in schema markup and AI search.
Adapting the template by page type
The six blocks hold across page types; what shifts is emphasis. On a definition page ("what is X"), block one carries almost the entire load: engines answering definitional questions lift the opening block nearly verbatim, so the 120 words deserve half the total writing time, and the evidence paragraphs exist mainly to make the definition trustworthy. On a comparison page ("X vs Y", "best X for Y"), the table is the star block, because the engine is being asked to compose exactly what the table contains; keep the verdict in prose above it, since Forrester's data shows buyers use assistants for vendor comparison more than any other commercial task, and an engine will happily carry a clear, attributed verdict.
On a pricing page or cost guide, block two dominates: concrete numbers with dates and billing terms are what engines extract, and stale prices are what they punish, so a visible updated date is part of the template. On a how-to or playbook, the question-shaped headings carry the structure, because procedural queries retrieve step passages; number the steps in the headings themselves. The FAQ block earns its slot on every type, and the answer-first block is never optional. If you build one template file for your team, parameterize emphasis by page type and keep the block order fixed.
A worked example: rebuilding a lost page
Suppose tracking shows you lose "best invoicing tool for freelance designers" and the incumbent page on your site is a feature tour titled "Invoicing, reimagined". Applying the template is a rebuild, and walking it block by block shows how the skeleton changes decisions. Block one: the new page opens with 110 words that answer the actual question, naming your product, its price, the two or three competitors a designer would shortlist, and the one number that differentiates you. That opening feels commercially reckless to marketers trained on teaser copy, and it is precisely what makes the passage liftable. Block two: every claim in the body gets dressed in evidence, so "designers love us" becomes a rating with a count and a platform name, and "saves time" becomes a measured minutes-per-invoice figure with the sample described.
Block three: a table comparing the shortlist on the columns designers ask about, price, proposals, contracts, payment rails, with your row told straight and rival rows told fairly, because a table an engine can verify against other sources is a table it can trust. Block four: the H2s become the sub-questions of the buying decision ("how much should a freelancer pay for invoicing?", "what happens when a client pays late?"). Block five: the FAQ absorbs the phrasings your sales inbox actually receives. Block six: title tag, description, canonical, Article and FAQPage schema, ten minutes, done. The rebuilt page is also simply a better page for humans, which is the quiet pattern in everything the GEO evidence rewards.
Common ways teams break the template
The teaser opening.The most frequent break: a first block that gestures at the answer ("choosing the right tool depends on many factors...") to protect a reveal further down. Engines do not scroll for payoffs. If block one does not commit, the page effectively has no block one.
Evidence laundering.Statistics with no named source, or sourced to "studies show". The Princeton lift came from citations an engine can attribute; a naked number reads as a claim, and claims lose to sourced numbers published by competitors. If the source cannot be named, the statistic should not ship.
The decorative table.A comparison table whose cells contain marketing adjectives ("powerful", "seamless") rather than values. Engines extract facts from tables; adjectives give them nothing to carry, and the block wastes its slot.
FAQ padding. Restating section headings as questions to hit a count. Engines deduplicate against the body, so a padded FAQ adds bulk without adding a single new retrievable passage. Fewer real questions beat more fake ones.
Shipping without indexing. The silent killer: a perfect page that never entered the indexes engines retrieve from. Publishing is step one of two; submission and confirmation of indexing is the step that makes the page exist. Then the recheck loop can begin.
From template to flipped answer
A template is a hypothesis until an engine confirms it. The full loop: pick a buying question you currently lose, build the page on these six blocks, get it indexed, then recheck the exact question repeatedly and compare mention rates before and after. Repeatedly is the operative word, because SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list, so one triumphant screenshot proves nothing in either direction.
This template is also, transparently, Reachroller's product spec. When Reachroller flags a question you lose, the 10-credit fix it generates arrives in exactly this shape: answer-first block, sourced evidence, table, FAQ, metadata, schema and indexing steps, and the recheck confirms whether the answer flipped. We publish the skeleton openly because the moat was never the structure; it is the loop that measures whether each page worked. Run it by hand with this article as the checklist, or let the tool run it from $29 per month.
Frequently asked questions
What is the most important block on a page AI engines cite?+
The answer-first block. Retrieval is passage-level: engines fetch chunks that look like complete answers, and a standalone 120-word block directly under the H1 is the chunk most likely to be lifted. If you change one thing about an existing page, move its conclusion to the top and make it quotable out of context.
How long should a citable page be?+
Long enough to answer one question completely, with evidence. Depth helps because it produces more extractable passages, but a 3,000-word page that gestures at twelve topics loses to a 1,500-word page that settles one. Structure beats length: every section should survive being read alone, because that is how engines read it.
Does schema markup make AI engines cite a page?+
The best current evidence says no, on its own. Ahrefs tracked 1,885 pages that added JSON-LD in 2025 and 2026 and measured citation changes within statistical noise across AI Mode, ChatGPT and AI Overviews. Add Article and FAQPage schema as parsing hygiene, then spend the real effort on visible evidence: statistics, named sources, quotations.
Why do tables help a page get cited?+
Because engines build comparisons, and a clean table is pre-structured comparison data. When an assistant composes a versus-style answer, rows and columns with plain values are far easier to extract accurately than prose. Tables also concentrate facts, and a caption naming the source gives the engine the attribution it prefers to repeat.
How is this template different from writing advice for AI citations?+
This is the skeleton; the writing craft is a separate skill. Our guide on writing content AI engines cite covers sentence-level decisions: how to phrase claims, what to cut, how honesty outperforms. This template covers load-bearing structure: which blocks exist, in what order, to what spec. Use the template to frame the page and the writing guide to fill it.
Should I retrofit old pages to this template or write new ones?+
Retrofit first. An indexed page with history and any existing authority converts to the template in a couple of hours: move the conclusion into an answer-first block, source the claims, convert the densest comparison into a table, add the FAQ. New pages make sense for questions you have no page for at all. Prioritize by tracking data, starting with lost questions where retrofit candidates already exist.
How do I know if a page built on this template worked?+
Recheck the exact question after the page is indexed, repeatedly, and compare mention rates before and after. Single checks are noise: SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. Reachroller automates the recheck and links every score to the raw answer, so a flip is a logged fact rather than an impression.
Sources referenced
- Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
- Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- 5W Research, ChatGPT citation share analysis, 2026
- Forrester, 2026 Buyers' Journey Survey (18,000 global business buyers)
- Botify, analysis of OpenAI crawl growth, 2026; OpenAI developer docs on OAI-SearchBot
Get the template written for you, then verified
Three days, 50 credits, every feature, no card. Track 25 questions and generate one fix page in exactly this shape, on your own domain.
Check my brand free