Google-Extended
Definition
Google-Extended is a robots.txt control token, introduced by Google in September 2023, that lets site owners decide whether their content may be used to train and ground Google's Gemini AI models. It is not a separate crawler and never appears in server logs; Googlebot does the crawling, and the token governs data use. Blocking it does not affect Google Search rankings or inclusion in AI Overviews.
What Google-Extended is, and what it is not
Google-Extended exists because Google reuses one crawl for many purposes. Googlebot fetches your pages once, and that crawl feeds the classic search index, AI Overviews and, unless you opt out, the training and grounding of Gemini models. Google introduced the Google-Extended token in September 2023 as the opt-out lever for the AI training part, after publishers objected that participating in search had silently become participating in model training.
The most common misunderstanding is expecting to see Google-Extended in server logs. You never will, because there is no Google-Extended crawler. It is a product token that Googlebot reads in your robots.txt and applies as a data-use rule downstream. Traffic keeps arriving under Googlebot's user agent either way, which is why log-based bot audits routinely miss whether a site has taken a position on it.
Per Google's documentation, disallowing Google-Extended tells Google not to use your content to improve Gemini models, covering both training and grounding, the mechanism by which Gemini retrieves live web content to support its answers. It is the narrowest of the major AI opt-outs: one company, one model family, one data use, with the crawl itself unaffected.
Google's crawler documentation defines the scope precisely: the token governs whether content may be used for Gemini Apps and for the Vertex AI generative APIs, including future generations of Gemini models. That scoping is why the decision keeps its shape over time. You are setting a standing policy on a model family rather than reacting to a single product, and Google applies the rule during crawl processing, so the same Googlebot visit serves search either way.
How to set Google-Extended in robots.txt
To opt out, add User-agent: Google-Extended followed by Disallow: / to your robots.txt. To opt out only part of the site, scope the path, for example Disallow: /research/. Doing nothing leaves you opted in, since Google treats the absence of a Google-Extended rule as permission to use crawled content for Gemini.
Because the token is applied by Googlebot, the rule takes effect through Google's normal crawl and processing cycle rather than instantly. There is nothing to verify in logs, and the only confirmation is your own robots.txt being correct and reachable. Test the file itself: a robots.txt that returns errors or sits behind a redirect chain can silently fail to register any of your directives, AI-related or otherwise.
Keep the tokens straight, because the failure mode is severe in one direction only. Blocking Google-Extended is a reversible data-use preference. Blocking Googlebot removes you from Google Search itself, along with AI Overviews and every surface built on the search index. Sites aiming a protest at Google's AI have occasionally hit the wrong token, and the second mistake costs actual traffic immediately. When in doubt, write the Google-Extended group explicitly and leave every Googlebot rule exactly as it was.
What blocking Google-Extended costs a brand
The direct cost today is small, which is exactly why the decision deserves care rather than reflex. Google states that Google-Extended does not affect Search rankings, and AI Overviews draw on the regular search index under Googlebot's rules, so a site that blocks Google-Extended keeps its rankings and keeps appearing in AI Overviews. What changes is Gemini's relationship to your content: your pages stop feeding Gemini training, and stop being available for grounding in Gemini products.
The grounding half is the part with commercial teeth. Grounding is how Gemini pulls live web content to support answers, and Gemini's answer surfaces keep expanding across the Gemini app and Google's AI products. Opting out of grounding means those answers lean on third-party descriptions of your brand instead of your own pages, with the usual consequences: staler facts, weaker framing, and citation opportunities forfeited to whoever remains in the pool.
The strategic logic mirrors every training opt-out. Large publishers with licensable archives block Google-Extended as negotiating leverage, and that position is coherent for them. For most brands, whose problem is being unknown rather than being exploited, contributing to the systems that answer buyers' questions is worth more than withholding. Wherever you land, treat it as one decision inside a full crawler policy covering OpenAI, Anthropic and Perplexity too, and verify the whole policy took effect with the free checker at /bot-access.
Frequently asked questions
Does blocking Google-Extended hurt my Google rankings?+
No. Google states that Google-Extended is not a ranking signal and that disallowing it does not affect your inclusion or position in Google Search. It also leaves AI Overviews eligibility untouched, since AI Overviews use the regular search index governed by Googlebot's rules. It only withdraws your content from Gemini training and grounding.
Why does Google-Extended never show up in my server logs?+
Because it is not a crawler. Googlebot performs the actual crawling under its own user agent, then applies your Google-Extended rule as a downstream data-use restriction. There is no separate fetch to observe, so the token is invisible in logs even when it is being honored.
Does blocking Google-Extended keep me out of AI Overviews?+
No. AI Overviews are built on Google's organic search index, and eligibility follows Googlebot and the standard search controls. Google-Extended governs Gemini training and grounding only. To stay out of AI Overviews you would need search-level controls like nosnippet directives, which carry their own costs in classic search.
Should a brand block Google-Extended?+
For most brands, no. The opt-out mainly serves publishers treating their archives as licensable assets. A brand trying to be recommended benefits from Gemini grounding on its own pages, since the alternative is Gemini describing the brand through third-party sources. Blocking is reversible, so it is a position you can revisit if licensing markets mature.
Sources referenced
- Google Search Central, Google crawlers and user agents documentation (Google-Extended product token)
- Google, The Keyword: An update on web publisher controls, September 2023
- Google Search Central documentation on AI Overviews and search snippet controls
See this metric on your own brand
Reachroller tracks the questions your buyers ask and shows exactly what AI answers. Three days free, no card.
Check my brand free