Glossary

PerplexityBot

Definition

PerplexityBot is the web crawler Perplexity operates to build the search index behind its AI answer engine. Perplexity documents that it surfaces and links websites in results and is not used to train AI models. Its user agent string is Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot), and it honors robots.txt rules addressed to the PerplexityBot token.

Who operates PerplexityBot and what it feeds

PerplexityBot is operated by Perplexity AI and documented in the company's crawler reference at docs.perplexity.ai. It feeds Perplexity's own search index, which the company has publicly described at more than 50 billion pages. Perplexity is an answer engine rather than a model vendor: it retrieves sources at question time and composes cited answers, so its index is the entire supply chain for what it can show users.

Perplexity states that PerplexityBot is not used for crawling content to train foundation models, which makes it a different kind of decision from GPTBot or ClaudeBot. Allowing PerplexityBot is closer to allowing Googlebot: it governs whether your pages can appear as linked, attributed sources in answers, rather than whether your content dissolves into a training corpus.

Perplexity also documents a second agent, Perplexity-User, with the user agent string Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user). It fetches pages when a user directly triggers a request, for example by pasting a URL. Perplexity states that because a human initiated the visit, Perplexity-User generally ignores robots.txt rules, a position that has drawn criticism from publishers and infrastructure providers.

How to allow or block PerplexityBot in robots.txt

PerplexityBot checks robots.txt and honors rules addressed to its token. To allow it explicitly, add User-agent: PerplexityBot followed by Allow: /. To block it, use User-agent: PerplexityBot followed by Disallow: /. Perplexity's own documentation recommends allowing the bot so your site can be indexed and surfaced in responses, and publishes the IP ranges it crawls from at perplexity.com/perplexitybot.json for verification.

Verification deserves extra weight with Perplexity specifically. In August 2025, Cloudflare published research alleging that Perplexity used undeclared crawlers with rotating user agents and IP ranges to reach content that had blocked its declared bots, which Perplexity disputed. Whatever position you take in that dispute, the practical lesson holds for all AI crawlers: robots.txt is a convention, user agent strings can be worn by anyone, and CDN-level bot management is the enforcement layer when access control genuinely matters.

Remember that Perplexity-User sits outside your robots.txt by Perplexity's own definition. If you need to stop user-triggered fetches as well, that requires blocking at the network level, at the cost of breaking every legitimate case where a prospective buyer asks Perplexity to read your pages.

What blocking PerplexityBot costs a brand

Perplexity is the engine where index access pays off most visibly, because it cites more sources per answer than any major competitor, roughly 8.2 on average by published measurements, several times more than ChatGPT. More citation slots per answer means more chances for a well-matched page from a smaller site to be shown, linked and attributed. A page PerplexityBot cannot reach is excluded from every one of those slots.

The cost compounds because of who uses Perplexity. Its user base skews toward research-heavy queries: people comparing products, checking claims and gathering sources before decisions. Those are bottom-of-funnel moments, and every answer in your category will cite somebody. Blocking the crawler does not make the question go away, it just guarantees the sources are your competitors and the framing of your market excludes you.

Citations also carry traffic and, for some publishers, revenue. Perplexity displays its sources prominently and links out, and in 2024 it launched a Publishers' Program that shares advertising revenue with participating media partners, a signal that the company treats source relationships as commercial infrastructure rather than raw material. Referral volumes from AI engines remain small next to classic search, but they land late in the buying journey, from users who asked a specific question and clicked the source behind the answer.

Since PerplexityBot does not feed model training, the usual argument for blocking training crawlers, withholding content as a licensing position, does not apply cleanly here. The trade is citation visibility against nothing concrete in return. Publishers with paywalled archives and licensing disputes have their own calculus, but for a brand whose economics depend on discovery, blocking PerplexityBot is a pure loss.

As with every crawler decision, confirm reality matches intent. CDN bot rules and security plugins frequently block PerplexityBot by default, and the free checker at /bot-access shows which AI crawlers can actually read your site. From there, the question becomes whether indexed pages turn into citations, which only checking real Perplexity answers can tell you.

Frequently asked questions

What is the difference between PerplexityBot and Perplexity-User?+

PerplexityBot is the index crawler: it browses the web on a schedule to build Perplexity's search index and honors robots.txt. Perplexity-User fetches a page only when a human directly triggers the request, and Perplexity states it generally ignores robots.txt for that reason. They have separate user agent strings and separate published IP ranges.

Does PerplexityBot train AI models on my content?+

Perplexity's documentation states PerplexityBot is used to surface and link websites in search results and is not used to crawl content for training AI foundation models. Its role is retrieval and citation, closer to a classic search crawler than to training bots like GPTBot or ClaudeBot.

How do I verify a real PerplexityBot visit?+

Match the requesting IP against the ranges Perplexity publishes at perplexity.com/perplexitybot.json, with a separate list for Perplexity-User. User agent strings are easily spoofed, and Perplexity's crawling has been the subject of public disputes, so IP verification is the only dependable way to attribute traffic before setting policy.

What happens to my Perplexity citations if I block PerplexityBot?+

Your pages drop out of the source pool Perplexity retrieves from, so they stop appearing among the roughly eight citations its average answer carries. Perplexity can still mention your brand through third-party pages that describe you, but you lose the ability to be a linked source in your own category's answers.

Keep reading

Sources referenced

  • Perplexity, Perplexity crawlers documentation (docs.perplexity.ai)
  • Perplexity, PerplexityBot IP ranges (perplexity.com/perplexitybot.json)
  • Cloudflare, research on undeclared crawling by Perplexity, August 2025

See this metric on your own brand

Reachroller tracks the questions your buyers ask and shows exactly what AI answers. Three days free, no card.

Check my brand free