SEO

PerplexityBot

Also called Perplexity crawler

The web crawler Perplexity uses to find and index pages it can cite in its AI-generated answers, controllable through robots.txt.

Quick facts: PerplexityBot

Category
SEO
Also called
Perplexity crawler
Level
Intermediate
Affects
Citations in Perplexity answers, server load, AI search visibility
Where to see it
robots.txt, server access logs, CDN or firewall bot settings, Perplexity's published IP ranges
In this article4
  1. How PerplexityBot works
  2. Why it matters for a UK business
  3. Common mistakes
  4. How to act on it

PerplexityBot is the web crawler used by Perplexity, the AI-powered answer engine, to find and index pages it can cite in its answers. It identifies itself with the user-agent token “PerplexityBot”, which means you can allow or block it in your robots.txt file like any other search crawler.

How PerplexityBot works

Perplexity answers questions by searching the web, reading the pages it finds and writing a summary with numbered links to its sources. PerplexityBot is the crawler that builds and refreshes the index those answers draw from. Perplexity’s own documentation says PerplexityBot is used to surface and link websites in its search results, and that it is not used to crawl content for training AI foundation models. It also publishes the IP address ranges the crawler uses, so a request claiming to be PerplexityBot can be checked against them.

There is a second agent to know about. When a Perplexity user asks it to visit a specific page, the fetch is made by an agent called Perplexity-User. Perplexity has said that because these fetches are made on a user’s request, they generally do not follow robots.txt rules. At the time of writing (October 2026), that is a meaningful difference from PerplexityBot, which does honour them. Perplexity’s crawling practices have also faced public criticism, including a 2025 report from Cloudflare alleging that it reached blocked sites through undeclared crawlers, which Perplexity disputed.

PerplexityBot is one of several AI crawlers alongside OpenAI’s GPTBot and OAI-SearchBot and Anthropic’s ClaudeBot. Each has its own user-agent and purpose, so each needs its own decision.

Why it matters for a UK business

More people now ask AI tools for recommendations and explanations rather than scrolling through ten blue links. If Perplexity cannot crawl your site, it cannot cite you, and a competitor’s page will be quoted instead. For a business that wants to be found as a source, such as a consultancy, a specialist retailer or a professional services firm publishing useful guides, letting PerplexityBot in is usually the sensible default.

Some publishers make the opposite choice. If your content is your product, for example paid research or a subscription publication, you may decide that being summarised without a click is a poor trade. That is a commercial decision, and robots.txt is the tool for expressing it. Blocking PerplexityBot has no effect on Google or Bing rankings.

Common mistakes

  • Blocking all bots with a blanket rule, or a security plugin, without realising AI search crawlers are caught too.
  • Assuming a robots.txt block also stops user-requested fetches from Perplexity-User.
  • Trusting the user-agent string alone. Anyone can claim to be PerplexityBot; check the IP against Perplexity’s published ranges before drawing conclusions from your logs.
  • Blocking the crawler, then wondering why the business is never mentioned in Perplexity’s answers.
  • Forgetting that a firewall or CDN bot setting can override what robots.txt says.

How to act on it

First decide what you want: to be cited by AI answer engines, or to keep your content out of them. Then check what actually happens. Read your robots.txt for rules naming PerplexityBot or applying to all agents, and look at your hosting or CDN settings, since some services now block AI crawlers by default unless you opt in. Your server logs will show whether PerplexityBot has visited and which pages it requested.

If you want visibility, allow the crawler, make sure important pages load without heavy JavaScript, and write content that states facts plainly so it can be quoted accurately. Then search Perplexity for the questions your customers ask and see who is cited. This sits within wider work on generative engine optimisation, which is what my AI search optimisation service covers.

Do and do not

Do

  • Decide deliberately whether to allow AI crawlers
  • Check CDN and firewall settings as well as robots.txt
  • Verify crawler IPs before trusting log entries

Do not

  • Block every bot with one blanket rule
  • Assume robots.txt stops user-requested fetches
  • Expect allowing it to change Google rankings

Questions people ask about this

How do I block PerplexityBot?

Add a group to your robots.txt file with "User-agent: PerplexityBot" followed by "Disallow: /". That tells the crawler not to fetch any page on the site. It does not cover Perplexity-User fetches made on a person's request, and it does not remove anything Perplexity indexed before the rule went in.

Does PerplexityBot use my content to train AI models?

Perplexity states that PerplexityBot is used for search and citations, not for training foundation models. That is the company's own description, and you cannot verify it from the outside. If your concern is training, review each AI company's documented crawlers separately, because they use different agents for search and for training.

Will allowing PerplexityBot help my Google rankings?

No. Google's rankings are not affected by whether you allow Perplexity's crawler. Allowing it only affects whether Perplexity can read and cite your pages in its own answers.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.