Playbooks
The GEO checklist: 30 checks before your next content sprint
Updated August 2, 2026
A GEO checklist is worth running before any content sprint because most AI visibility failures happen outside the writing: a blocked crawler, a fact trapped in an image, a review profile nobody claimed, a score computed from a single lucky run. This checklist covers 30 checks in four groups, crawlability, content structure, third-party presence and measurement, each grounded in published evidence, from the Princeton GEO study's finding that statistics, quotations and cited sources lift generative visibility by up to 40 percent to SparkToro's demonstration that identical ChatGPT runs almost never return the same brand list. Work the groups in order; each gates the next. Reachroller automates the measurement group, and its receipts tell you which of the other checks to fix first.
How to use this checklist
Content sprints fail quietly in GEO. A team writes six good pages, waits a month, sees no movement in AI answers and concludes the discipline is hype, when the actual culprit was a firewall rule from 2024 or a score computed from one lucky ChatGPT run. The checklist exists to catch those failures before the writing starts. It is ordered by dependency: crawlability gates content, content gates citation, third-party presence corroborates it all, and measurement is how you find out which of the first three is your bottleneck.
Every check cites its evidence, most heavily the Princeton GEO study, which tested nine optimization methods across 10,000 queries and found statistics, quotations and cited sources lifting visibility in generative answers by up to 40 percent while keyword stuffing fell below baseline. If the discipline itself is new to you, read what is generative engine optimization first, and GEO vs SEO for how these checks relate to the SEO audit you already know. Then work the groups in order and mark each check pass, fail or unknown; the unknowns are usually where the answers went.
Crawlability: can engines fetch you at all?
Everything downstream depends on retrieval. AI engines pull sources from search indexes and their own crawlers, so a page they cannot fetch or parse does not exist for them, however brilliant the writing. These eight checks fail silently and often; the fix is usually minutes once found.
| # | Check | Why it matters |
|---|---|---|
| 1 | robots.txt allows answer-engine crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) | Search crawlers make you eligible for citations. OpenAI's OAI-SearchBot has roughly tripled its crawl since August 2025 per Botify; blocking it removes you from the fastest-growing answer surface. |
| 2 | Training and search crawlers handled as separate decisions | GPTBot is the most-blocked AI crawler in robots.txt analyses. Blocking training bots is a legitimate policy choice; copying a blocklist that also bans search bots silently deletes your citation eligibility. |
| 3 | Indexed in Google AND Bing, key URLs submitted in both consoles | Google's AI features cite from Google's index; the ChatGPT ecosystem grew up on Bing's. A page missing from either index is invisible to the engines that retrieve from it. |
| 4 | XML sitemap current and referenced in robots.txt | Retrieval-grounded answers favor fresh fetches. A stale sitemap slows discovery of exactly the fix pages your sprint will publish, delaying the recheck that proves they worked. |
| 5 | Key content renders as server-side HTML text, never JS-only or image-only | Most AI crawlers read raw HTML. Pricing in a JavaScript widget or specs in a JPEG do not exist for retrieval, so the engine quotes a competitor whose facts are in text. |
| 6 | CDN, WAF and bot protection verified against AI crawler user agents in logs | AI bot traffic grew over 300 percent between January 2025 and March 2026, and default firewall rules often challenge or block it silently. Your logs are the only place this failure is visible. |
| 7 | Fast, stable pages with no interstitials gating first content | Crawlers work on budgets. Slow responses, consent walls and popups that displace content reduce what gets fetched and parsed, which reduces what can be cited. |
| 8 | Clean canonicals, one URL per answer, no duplicate confusion | Engines resolve entities and sources. Three near-identical URLs split your retrieval signal and give the engine a reason to trust an aggregator's single clean page instead. |
Content structure: is there anything worth quoting?
Retrieval puts your page on the engine's desk; structure decides whether the answer uses it. The Princeton GEO study is the evidence base here: statistics, quotations and cited sources lifted visibility by up to 40 percent while keyword stuffing fell below baseline. These eight checks turn that finding into an editing pass.
| # | Check | Why it matters |
|---|---|---|
| 9 | Every key page opens with an answer-first block of roughly 100 to 120 words | Engines extract passages, and a complete, quotable answer under the H1 is the easiest passage to lift. Pages that wind up to their point lose to pages that lead with it. |
| 10 | Statistics with named sources appear throughout | The Princeton GEO study measured statistics among the strongest visibility lifts, up to 40 percent in generative answers. Unsourced numbers read as marketing; sourced ones read as evidence. |
| 11 | Quotations from named people are present and attributable | Quotation addition was a top-performing method in the same Princeton research. Named, checkable voices give engines human authority to relay, which anonymous prose cannot supply. |
| 12 | Claims cite outbound sources you actually link | Cited sources were the third top-performing Princeton method. Pages that reference evidence behave like the sources engines already trust, and get treated accordingly. |
| 13 | Zero keyword stuffing, phrasing written for a reader | Princeton measured keyword stuffing below baseline for generative engines, worse than changing nothing. The oldest SEO reflex is now a measurable liability. |
| 14 | One page answers one question; H2s and H3s map to sub-questions | Retrieval matches questions to passages. A page that answers one buying question completely beats a page that gestures at twelve, and a question-shaped heading is a retrieval anchor. |
| 15 | Comparison tables and structured facts exist in HTML, not prose alone | Engines extract structured claims from tables more reliably than from paragraphs. Prices, features and trade-offs in table rows become the skeleton of composed comparisons. |
| 16 | Schema markup present: Article, FAQPage, Product or LocalBusiness as applicable | Ahrefs' May 2026 study of 1,885 pages found schema is no magic citation switch, but it disambiguates entities, prices and questions cheaply. Treat it as hygiene that removes excuses, never as strategy. |
Third-party presence: who vouches for you?
Engines compose verdicts from sources you do not own: 5W Research puts Wikipedia and Reddit alone at over a quarter of U.S. ChatGPT citations. Half of GEO is an away game, and these seven checks are its fixture list.
| # | Check | Why it matters |
|---|---|---|
| 17 | Review-site profiles claimed, complete and recently reviewed | G2 reports AI chatbots now shape 54 percent of software shortlists, and engines corroborate vendors through review platforms. An unclaimed or stale profile is a missing character witness. |
| 18 | Wikipedia and Wikidata presence accurate where your brand merits entries | 5W Research measured Wikipedia at 13.15 percent of U.S. ChatGPT citations, the single largest source. You rarely control it, but factual errors there propagate into answers. |
| 19 | Reddit and community threads about your category monitored, participated in honestly | Reddit carries 11.97 percent of U.S. ChatGPT citations and more on Perplexity. Disclosed, substantive participation builds citable text; astroturf gets archived and quoted against you. |
| 20 | The listicles and rankings engines cite for your category identified, inclusion pitched | Best-of questions retrieve pages that already rank for best-of phrases, which are third-party lists. The citation trail on questions you lose names exactly which lists matter. |
| 21 | Original, quotable data published for others to cite | Firms and brands that publish numbers become the source other pages quote, and engines cite the cited. One real benchmark outperforms ten opinion posts. |
| 22 | Entity facts consistent everywhere: name, category, pricing, locations | Engines cross-reference sources before naming you. Contradictions between your site, profiles and directories read as uncertainty, and confident answers route around uncertain entities. |
| 23 | A per-engine source map maintained from actual answer citations | Only about 11 percent of cited domains overlap between ChatGPT and Perplexity in cross-platform analyses. The away game differs per engine, so the map must too. |
Measurement: would you know if any of this worked?
AI answers are probabilistic, so measurement discipline separates GEO programs from GEO theater. These seven checks define the honest loop: stable questions, repeated runs, receipts, trends, and a recheck after every fix.
| # | Check | Why it matters |
|---|---|---|
| 24 | A fixed list of unbranded buying questions, phrased conversationally | You cannot trend what you keep rephrasing. Fifteen to twenty-five questions in your buyers' own words, held stable across weeks, is the instrument everything else reads from. |
| 25 | Repeated runs on a schedule, never single checks | SparkToro measured under a 1 percent chance that two identical ChatGPT runs return the same brand list. One check is a coin flip with a screenshot. |
| 26 | Branded questions excluded from the headline score | A question containing your brand name mentions you by construction. Mixing those in inflates the score and hides the losses that matter. Reachroller excludes them by policy. |
| 27 | Raw answers and citations stored as auditable receipts | A score without the answer behind it cannot be debugged or believed. Receipts turn a dashboard number into a work queue: which question, which engine, which sources. |
| 28 | Mention rate tracked per question and per engine, trended over weeks | Engines disagree and answers drift. Per-question, per-engine trend lines are the only honest signal of progress; averages across engines hide exactly the gaps you need. |
| 29 | AI referral traffic segmented in analytics | Referrals from chatgpt.com, perplexity.ai and gemini.google.com convert differently, 31 percent better than non-branded organic in one 94-site ecommerce analysis. A separate segment makes GEO a channel with revenue attached. |
| 30 | Every published fix followed by a recheck against baseline | The loop closes or the sprint was theater. Publish, index, recheck, compare mention rates. A fix that does not move the trend goes back in the queue with a new hypothesis. |
Reading your results: three common profiles
Teams that run the audit tend to land in one of three profiles. The blocked publisher passes the content checks and fails crawlability: strong pages, invisible to engines, usually via a bot-protection default or a robots.txt copied from a paranoid template. The fix is fast and the payoff is fast, because the content already exists. The self-referential brand passes crawlability and content but fails the third-party group: everything the engines could say about them is something the brand said about itself, so cautious engines say nothing. That profile needs the slow work of reviews, communities and earned citations, started now because it compounds.
The optimist is the most common profile: reasonable passes across the first three groups and a measurement section full of unknowns. They checked ChatGPT once, saw their name, and stopped looking. Given SparkToro's finding that identical runs agree less than 1 percent of the time, that single check carries almost no information. The optimist does not need more content; they need a baseline, which is the cheapest item on the entire checklist to fix.
Whichever profile you match, resist the urge to fix everything at once. The checklist is a diagnosis; the treatment is one bottleneck at a time, verified by the recheck in check 30, because a loop that verifies beats a sprint that assumes.
What this checklist deliberately leaves out
Some fashionable items are absent on purpose. llms.txt is the loudest: large-scale 2026 studies spanning hundreds of thousands of domains and hundreds of millions of bot events found no measurable citation lift from the file by itself, so it earns a footnote rather than a check. Add one if you like, it costs minutes, but never let it substitute for the crawl access and structure checks that actually gate retrieval. Prompt-injection tricks, hidden text addressed to the models, sit further down the list still: they are detectable, embarrassing when archived, and aimed at systems that update faster than the trick spreads.
Also absent: anything you would buy instead of earn. Paid placement in answer engines barely exists as of mid-2026, and the citation economics described in check 20 reward the slower route of being genuinely present in the sources engines already trust. Incentivized reviews and undisclosed community seeding belong in the same bin, since platforms prosecute both and engines quote the prosecution. The checklist is boring by design, because every check on it survives an engine update, an algorithm leak and a journalist reading your robots.txt.
Automate the fourth group, then start the sprint
Groups one through three are human work: log reading, editing, profile claiming, honest participation. Group four is machine work, and doing it by hand is how measurement quietly stops happening. Reachroller runs the entire measurement group as designed here: a fixed question list run on schedule across engines, mention rates with branded questions excluded, every raw answer and its citations stored as receipts, trends per question and per engine, and a recheck workflow after each fix. When a question is lost, it generates the fix page publish-ready, with slug, title tag, meta description and schema, which turns checks 9 through 16 into a template instead of a memory. Starter is $29 per month for 400 credits and 25 tracked questions; the honest caveat is that ChatGPT tracking is live today and the remaining engines are rolling out. How that stacks against other trackers is covered in the best AI visibility tools.
The sequence for your next sprint, then: run the audit, fix crawlability the same day, baseline your questions with a Reachroller check, and only then write, with checks 9 through 16 taped above the keyboard. Four weeks later the recheck tells you what the sprint actually earned, in mention rates rather than vibes. The free check gives you three days, 50 credits and every feature, no card, which is enough to complete the measurement group of this checklist before you write a single new page.
Frequently asked questions
How long does the full 30-check audit take?+
For a typical small site, one focused day. Crawlability checks are an hour with server logs and a robots.txt read. Content structure is a page-by-page editing pass on your ten most important URLs. Third-party presence takes an afternoon of profile claims and citation reading. Measurement setup is fastest with tooling: Reachroller's free check baselines your questions in minutes.
Which group matters most if I can only fix one?+
Order matters more than weight: crawlability gates everything, so a blocked crawler makes the other 22 checks irrelevant. Once retrieval works, content structure carries the strongest published evidence, the Princeton study's up-to-40-percent lift. But if you already publish decent content, the measurement group usually changes behavior fastest, because receipts reveal which specific checks are costing you answers.
Do I need an llms.txt file?+
It is optional and unproven. Large-scale 2026 studies across hundreds of thousands of domains found no measurable citation lift from llms.txt by itself, which is why it appears nowhere in these 30 checks. It costs little and may help agents navigate large docs sites, but treat it as an experiment, never as a substitute for crawl access, structure or third-party evidence.
How often should I rerun the checklist?+
Crawlability and third-party checks quarterly, or after any site migration, firewall change or redesign, which is when silent breakage happens. Content structure checks belong in your publishing workflow, applied to every new page. Measurement runs continuously by definition: weekly scheduled runs with a monthly review of trends against the fixes you shipped.
Does schema markup guarantee AI citations?+
No. Ahrefs' May 2026 analysis of 1,885 pages found no strong causal lift from schema alone, and this checklist ranks it as hygiene rather than strategy. Schema removes ambiguity about entities, prices and questions cheaply, which prevents certain failures. What earns citations is the combination the rest of the checklist builds: retrievable pages, quotable evidence, third-party corroboration.
Can one person run this whole loop?+
Yes, if the tracking and drafting are automated. The checklist was sized for a founder or a one-person marketing team: crawlability is a one-time fix, structure is a template, third-party work is a steady trickle, and measurement is the part machines do better anyway. Reachroller Starter runs the measurement group and generates fix pages for $29 per month, which is the entire tooling budget the loop requires.
Sources referenced
- Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
- SparkToro, consistency of repeated ChatGPT brand recommendations, 2025
- 5W Research, ChatGPT citation share analysis, 2026
- Botify, analysis of OpenAI crawl growth, 2026
- Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
- 2026 robots.txt and AI crawler traffic analyses (crawler blocking rates, AI bot traffic growth)
- 2026 large-scale llms.txt studies across 300,000+ domains
- G2, Buyer Behavior Report, March 2026
- Analysis of 94 ecommerce sites, ChatGPT referral conversion, reported by Search Engine Land, 2026
Checks 24 through 30, done for you by tonight
Three days, 50 credits, every feature, no card. Baseline your buying questions before the sprint, so the recheck means something.
Check my brand free