GPTBot is the web crawler OpenAI uses to collect publicly available pages that may be used to train its AI models, the systems behind ChatGPT. It identifies itself with the user agent token GPTBot, and site owners can allow or block it in their robots.txt file.
How GPTBot works
Like Googlebot, GPTBot requests pages, follows links and stores what it finds. The difference is the purpose: Googlebot builds a search index, while GPTBot gathers material that may help train future large language models.
OpenAI runs more than one crawler, and each has its own token. According to OpenAI’s documentation at the time of writing (October 2026):
- GPTBot collects content for model training.
- OAI-SearchBot finds and indexes pages so they can appear as sources in ChatGPT’s search answers.
- ChatGPT-User fetches a page when a person using ChatGPT asks it to visit a link. Because a person starts these visits, OpenAI handles them differently from automated crawling, so check its current documentation.
Blocking takes two lines in robots.txt: a “User-agent: GPTBot” line followed by “Disallow: /”. You can also block only certain folders. OpenAI publishes the IP ranges its crawlers use, so a firewall can confirm that a visitor claiming to be GPTBot really is.
Why it matters
The choice has commercial consequences. Blocking GPTBot means your future content is less likely to shape what ChatGPT knows about your subject, your brand or your prices. Allowing it means your writing may be used to train a commercial product without payment or credit. Neither is wrong; it depends on what you publish.
A publisher whose income depends on original articles may decide the trade is a bad one. A plumbing firm in Leicester or a dental practice in Glasgow usually wants AI assistants to describe its services accurately, and has little to lose from being read.
The key point is that blocking GPTBot does not, on OpenAI’s account, remove you from ChatGPT search results. That depends on OAI-SearchBot. Blocking both cuts you off from being cited in ChatGPT’s answers altogether.
In the UK, policy on AI training and copyright was still being worked out at the time of writing, with the Government consulting on how text and data mining should be treated. So blocking is a business decision for now, not something the law requires either way.
Common mistakes
- Blocking every AI crawler without separating training from search. You can block training and still allow search crawlers.
- Not knowing your host or CDN blocks AI bots. Some providers, including Cloudflare, offer AI crawler blocking, and on some accounts it is on by default.
- Assuming robots.txt removes content already collected. It only affects future crawling.
- Copying a robots.txt from another site. A careless rule can block Googlebot as well.
How to act on it
Decide your position for training and for AI search separately, then write it down so it survives your next website change. Check your current robots.txt at yourdomain.co.uk/robots.txt, and look at your CDN or security plugin settings for AI crawler controls. Your server logs will show whether GPTBot is visiting and how often.
If you want to appear in AI answers, allow the search crawlers and make sure your key pages are crawlable. Working out which AI crawlers to allow, and making a site easy for them to read and cite, is part of my AI search optimisation work.
