PerplexityBot is the web crawler used by Perplexity, the AI-powered answer engine, to find and index pages it can cite in its answers. It identifies itself with the user-agent token “PerplexityBot”, which means you can allow or block it in your robots.txt file like any other search crawler.
How PerplexityBot works
Perplexity answers questions by searching the web, reading the pages it finds and writing a summary with numbered links to its sources. PerplexityBot is the crawler that builds and refreshes the index those answers draw from. Perplexity’s own documentation says PerplexityBot is used to surface and link websites in its search results, and that it is not used to crawl content for training AI foundation models. It also publishes the IP address ranges the crawler uses, so a request claiming to be PerplexityBot can be checked against them.
There is a second agent to know about. When a Perplexity user asks it to visit a specific page, the fetch is made by an agent called Perplexity-User. Perplexity has said that because these fetches are made on a user’s request, they generally do not follow robots.txt rules. At the time of writing (October 2026), that is a meaningful difference from PerplexityBot, which does honour them. Perplexity’s crawling practices have also faced public criticism, including a 2025 report from Cloudflare alleging that it reached blocked sites through undeclared crawlers, which Perplexity disputed.
PerplexityBot is one of several AI crawlers alongside OpenAI’s GPTBot and OAI-SearchBot and Anthropic’s ClaudeBot. Each has its own user-agent and purpose, so each needs its own decision.
Why it matters for a UK business
More people now ask AI tools for recommendations and explanations rather than scrolling through ten blue links. If Perplexity cannot crawl your site, it cannot cite you, and a competitor’s page will be quoted instead. For a business that wants to be found as a source, such as a consultancy, a specialist retailer or a professional services firm publishing useful guides, letting PerplexityBot in is usually the sensible default.
Some publishers make the opposite choice. If your content is your product, for example paid research or a subscription publication, you may decide that being summarised without a click is a poor trade. That is a commercial decision, and robots.txt is the tool for expressing it. Blocking PerplexityBot has no effect on Google or Bing rankings.
Common mistakes
- Blocking all bots with a blanket rule, or a security plugin, without realising AI search crawlers are caught too.
- Assuming a robots.txt block also stops user-requested fetches from Perplexity-User.
- Trusting the user-agent string alone. Anyone can claim to be PerplexityBot; check the IP against Perplexity’s published ranges before drawing conclusions from your logs.
- Blocking the crawler, then wondering why the business is never mentioned in Perplexity’s answers.
- Forgetting that a firewall or CDN bot setting can override what robots.txt says.
How to act on it
First decide what you want: to be cited by AI answer engines, or to keep your content out of them. Then check what actually happens. Read your robots.txt for rules naming PerplexityBot or applying to all agents, and look at your hosting or CDN settings, since some services now block AI crawlers by default unless you opt in. Your server logs will show whether PerplexityBot has visited and which pages it requested.
If you want visibility, allow the crawler, make sure important pages load without heavy JavaScript, and write content that states facts plainly so it can be quoted accurately. Then search Perplexity for the questions your customers ask and see who is cited. This sits within wider work on generative engine optimisation, which is what my AI search optimisation service covers.
