Google-Extended is a name you can use in your robots.txt file to tell Google not to use your site’s content to train its Gemini AI models or to ground Gemini’s answers. It is not a separate crawler; it is a control token that changes what Google does with pages its existing crawlers have already fetched, and it has no effect on whether your site appears in Google Search.
How Google-Extended works
Google introduced the token in 2023. You add it to robots.txt like any other user agent group, for example a line reading User-agent: Google-Extended followed by Disallow: / to opt the whole site out, or Disallow lines for specific folders only.
Because no crawler announces itself as Google-Extended, you will not find it in your server logs. Pages are still fetched by Googlebot and Google’s other crawlers; the token simply records your choice about how that content may be used. According to Google’s crawler documentation at the time of writing (October 2026), the opt-out covers use of your content for training future Gemini models and for grounding answers in Gemini Apps and the Gemini API on Vertex AI.
The boundary that catches most people out is search. Google’s AI features within Search, including AI Overviews, are governed by Googlebot, not by Google-Extended. Blocking Google-Extended does not remove your pages from AI Overviews, and the only way to do that through robots.txt would be to block Googlebot, which removes you from Google Search altogether. Other controls, such as nosnippet, limit what can be shown from a page, but they also affect ordinary search snippets. Google has changed these controls before, so check the current documentation before deciding.
Why it matters
For a UK business, this is a commercial decision rather than a technical one. If your content is your product, such as original research, training material, recipes or detailed guides, you may not want it used to build a product that answers questions without sending visitors to you. If your content exists mainly to win customers, being drawn on and cited by AI assistants may be a source of visibility you want.
The legal background is also unsettled. The UK government has consulted on how copyright law should apply to AI training, and at the time of writing the rules had not been finalised. A robots.txt token is not a legal instrument, but it is a clear, public record of your preference, and it is easy to change if your view or the law changes.
Common mistakes
- Thinking it blocks AI Overviews. It does not; those sit within Google Search and follow Googlebot rules.
- Blocking Googlebot by accident. A rule meant for AI crawlers that is written under User-agent: * or Googlebot will remove the site from search.
- Assuming it covers every AI company. It applies only to Google. OpenAI’s GPTBot, Anthropic’s ClaudeBot and others each have their own user agent names.
- Looking for it in log files. It never appears there, because it is not a crawler.
- Deciding without agreement. Marketing may want AI visibility while publishing teams want protection; settle the policy first.
How to act on it
Decide your position on AI use of your content, ideally per section of the site: you might allow marketing pages while opting out a paid resource library. Then open your robots.txt (yoursite.co.uk/robots.txt) and see what it says now. Some plugins, hosting services and content delivery networks can add AI crawler blocks on your behalf, so check whether a choice has already been made for you.
If you change it, add a separate group for Google-Extended rather than editing the Googlebot rules, and test the file in Google Search Console’s robots.txt report afterwards. Review the decision whenever Google updates its documentation. Working out where your content should and should not appear in AI answers is part of the AI search optimisation work I do, alongside understanding how grounding decides which sources get cited.
