OAI-SearchBot is the web crawler OpenAI uses to find and index pages so they can be shown and linked as sources in ChatGPT’s search features. If it cannot reach your site, your pages are far less likely to appear as cited results when someone uses ChatGPT to look something up.
How OAI-SearchBot works
At the time of writing (October 2026), OpenAI runs several separate user agents, each with its own job:
- OAI-SearchBot crawls for search, building the index ChatGPT draws on when it looks things up and links to sources.
- GPTBot collects content that may be used to train OpenAI’s models.
- ChatGPT-User fetches a page when a person’s request leads ChatGPT to visit it directly. It acts for that person rather than crawling the web, and OpenAI says robots.txt rules may not apply to these user-initiated visits.
Because OAI-SearchBot and GPTBot are controlled separately in robots.txt, you can allow one and block the other. A common choice is to let OAI-SearchBot in, so the site can be cited, while disallowing GPTBot, so the content is not used for training. OpenAI publishes the IP ranges its crawlers use, which lets you confirm that a visitor calling itself OAI-SearchBot genuinely is, since user-agent strings are easy to fake.
OpenAI’s documentation says robots.txt changes can take around a day to be reflected. It also says a site that blocks OAI-SearchBot may still appear as a plain navigational link, but its content will not be used in ChatGPT’s search answers.
Why it matters
A growing number of people begin research in AI assistants rather than a search box. When ChatGPT answers a question such as “which conveyancing solicitors in Leeds have good reviews”, it cites the pages it relied on, and those citations bring visits and credibility. A site that blocks OAI-SearchBot, often without meaning to, removes itself from that pool.
Accidental blocking is common. Some security plugins, CDN bot settings and hosting firewalls turn away unfamiliar crawlers by default, and some owners added blanket AI blocks in 2023 without separating training crawlers from search crawlers. The result is a business that would happily be recommended by ChatGPT but has told its search crawler to keep out.
For a UK business, this is a commercial decision about whether visibility in AI answers is worth more than any concern about how the content is used. Separate user agents mean you can decide each part on its own merits.
Common mistakes
- One rule for every AI crawler. Blocking all of them removes search visibility along with training access. The same thinking applies to other assistants’ crawlers, such as PerplexityBot.
- Checking robots.txt only. A firewall or CDN rule can return errors to OAI-SearchBot even when robots.txt allows it.
- Trusting the user agent alone. Scrapers impersonate well-known bots, so verify against OpenAI’s published IP ranges before giving any special access.
- Expecting access to guarantee citations. Being crawlable makes a page eligible. Whether it is cited depends on how clearly and usefully it answers the question.
How to act on it
Open your robots.txt file and look for rules naming OAI-SearchBot, GPTBot, or a wildcard that blocks all crawlers. Then check your server or CDN logs for requests from OAI-SearchBot and the status codes they received; a run of 403 errors means something upstream is turning it away. Log file analysis is the reliable way to see what any crawler actually does on your site.
Decide your policy for each OpenAI agent and write it into robots.txt explicitly, so the next developer understands the intent. Then make sure the pages you most want cited answer specific questions plainly, with the business name, location and key facts easy to extract. That combination of access and content is what my generative engine optimisation work covers, and the general rules for crawler access are set out in the robots.txt entry.
