ClaudeBot is the web crawler run by Anthropic, the company that makes the Claude AI assistant. It visits publicly available web pages and collects their content, which Anthropic says may be used to develop and train its models. It identifies itself with the user agent name “ClaudeBot” and, like other well-behaved crawlers, reads your robots.txt file before fetching pages.
How ClaudeBot works
A crawler requests a page, stores what comes back and follows links to find more. ClaudeBot does this at scale across the public web. It does not log in, fill in forms or reach pages behind a paywall or password.
At the time of writing (October 2026), Anthropic documents more than one bot, each with a separate purpose and a separate user agent:
- ClaudeBot collects content that may contribute to model training.
- Claude-User fetches a page when a person using Claude asks it to look at that page or answer a question that needs it.
- Claude-SearchBot crawls to improve the quality of search results shown to Claude users.
Because each has its own name, a rule in robots.txt for one does not apply to the others. Anthropic states that its bots respect robots.txt, and it lists the details on its own support pages, which are worth checking because names and behaviour in this area change. The arrangement is similar to OpenAI’s, which separates GPTBot for training from its search crawler.
Like many AI crawlers, ClaudeBot appears to work mainly from the HTML your server sends rather than running every script on the page, so content that only appears after JavaScript loads may not be collected.
Why it matters
For a UK business, the question is whether you want your public content available to AI systems. There are reasonable views on both sides. Allowing these bots means your service descriptions, prices and guides can inform the answers Claude gives, and the search and user-fetch bots in particular are how a page can be found and cited when someone asks Claude a question. Blocking the training crawler keeps your content out of future training data, which some publishers and creative businesses prefer.
Two practical points often get missed. Blocking ClaudeBot does not stop Claude from fetching a page when a user specifically asks for it unless you also address Claude-User. And blocking the search crawler may reduce the chance of being cited in Claude’s answers. Decide per bot, according to what you want, rather than copying a list from another site.
Crawl volume can matter too. On small or slow hosting, heavy crawling by several AI bots at once can affect performance, which shows up in server logs as bot traffic.
Common mistakes
- Blocking everything by accident. A blanket “Disallow: /” for all user agents, left over from a staging site, blocks Google as well as AI bots.
- Assuming one rule covers all Anthropic bots. Each user agent needs its own entry if you want different behaviour.
- Trusting the user agent alone. Anyone can send a request claiming to be ClaudeBot. Unusual traffic under that name is not proof it came from Anthropic.
- Expecting robots.txt to remove past data. It controls future crawling, not content already collected.
How to act on it
Look at your current robots.txt (it lives at yourdomain.co.uk/robots.txt) and check whether it mentions ClaudeBot, Claude-User or Claude-SearchBot. Then decide your policy: allow all three, block only the training crawler, or block all of them. To block only training, add a group with “User-agent: ClaudeBot” followed by “Disallow: /”, and leave the others alone.
Check your server logs or hosting dashboard for requests from these user agents to see how often they visit. Make sure your important content is in the HTML rather than loaded by scripts. If you want your business to be found and cited in AI answers, this decision belongs in a wider plan, which is what my AI search optimisation work covers.
