SEO

GPTBot

Also called OpenAI crawler

OpenAI's web crawler that gathers public pages which may be used to train its AI models. You control it through robots.txt.

Quick facts: GPTBot

Category
SEO
Also called
OpenAI crawler
Level
Intermediate
Affects
Whether your content is used in AI training, server load, AI visibility decisions
Where to see it
robots.txt, server access logs, CDN bot settings such as Cloudflare, OpenAI's published crawler documentation
In this article4
  1. How GPTBot works
  2. Why it matters
  3. Common mistakes
  4. How to act on it

GPTBot is the web crawler OpenAI uses to collect publicly available pages that may be used to train its AI models, the systems behind ChatGPT. It identifies itself with the user agent token GPTBot, and site owners can allow or block it in their robots.txt file.

How GPTBot works

Like Googlebot, GPTBot requests pages, follows links and stores what it finds. The difference is the purpose: Googlebot builds a search index, while GPTBot gathers material that may help train future large language models.

OpenAI runs more than one crawler, and each has its own token. According to OpenAI’s documentation at the time of writing (October 2026):

  • GPTBot collects content for model training.
  • OAI-SearchBot finds and indexes pages so they can appear as sources in ChatGPT’s search answers.
  • ChatGPT-User fetches a page when a person using ChatGPT asks it to visit a link. Because a person starts these visits, OpenAI handles them differently from automated crawling, so check its current documentation.

Blocking takes two lines in robots.txt: a “User-agent: GPTBot” line followed by “Disallow: /”. You can also block only certain folders. OpenAI publishes the IP ranges its crawlers use, so a firewall can confirm that a visitor claiming to be GPTBot really is.

Why it matters

The choice has commercial consequences. Blocking GPTBot means your future content is less likely to shape what ChatGPT knows about your subject, your brand or your prices. Allowing it means your writing may be used to train a commercial product without payment or credit. Neither is wrong; it depends on what you publish.

A publisher whose income depends on original articles may decide the trade is a bad one. A plumbing firm in Leicester or a dental practice in Glasgow usually wants AI assistants to describe its services accurately, and has little to lose from being read.

The key point is that blocking GPTBot does not, on OpenAI’s account, remove you from ChatGPT search results. That depends on OAI-SearchBot. Blocking both cuts you off from being cited in ChatGPT’s answers altogether.

In the UK, policy on AI training and copyright was still being worked out at the time of writing, with the Government consulting on how text and data mining should be treated. So blocking is a business decision for now, not something the law requires either way.

Common mistakes

  • Blocking every AI crawler without separating training from search. You can block training and still allow search crawlers.
  • Not knowing your host or CDN blocks AI bots. Some providers, including Cloudflare, offer AI crawler blocking, and on some accounts it is on by default.
  • Assuming robots.txt removes content already collected. It only affects future crawling.
  • Copying a robots.txt from another site. A careless rule can block Googlebot as well.

How to act on it

Decide your position for training and for AI search separately, then write it down so it survives your next website change. Check your current robots.txt at yourdomain.co.uk/robots.txt, and look at your CDN or security plugin settings for AI crawler controls. Your server logs will show whether GPTBot is visiting and how often.

If you want to appear in AI answers, allow the search crawlers and make sure your key pages are crawlable. Working out which AI crawlers to allow, and making a site easy for them to read and cite, is part of my AI search optimisation work.

Do and do not

Do

  • Decide separately whether to allow AI training and AI search crawlers
  • Check your CDN and security plugin for AI bot blocking you did not choose
  • Test robots.txt changes so Googlebot is unaffected

Do not

  • Block OAI-SearchBot by accident when you only meant to block training
  • Assume a robots.txt rule removes content already collected
  • Treat blocking as a UK legal requirement

Questions people ask about this

Will blocking GPTBot stop my site appearing in ChatGPT?

Not according to OpenAI. GPTBot gathers content for model training, while OAI-SearchBot handles the pages ChatGPT can show and cite in search answers. If you block GPTBot but allow OAI-SearchBot, your pages can still appear as sources. Block both and you lose that visibility.

Does blocking GPTBot affect my Google rankings?

No. Google ranks pages using Googlebot, which is a separate crawler with its own robots.txt rules. A GPTBot rule has no effect on Google as long as it is written correctly. Test any robots.txt change to make sure you have not blocked Googlebot by mistake.

Can GPTBot ignore my robots.txt?

OpenAI says GPTBot respects robots.txt, and in practice well-known AI crawlers generally follow it. Robots.txt is a convention rather than an enforcement tool, though, and anyone can fake a user agent name. If you need certainty, check the requests against OpenAI's published IP ranges and block at the firewall.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.