PLATFORMS

How Perplexity Decides What to Cite

IN ONE SENTENCE

Getting cited in Perplexity requires content to be inside its retrievable set and, for a given question, more worth citing than the alternatives.

Being named in Perplexity requires two things: the content is inside its retrievable set, and for a given question it is more worth citing than the alternatives. Perplexity states that PerplexityBot 'indexes pages similarly to other search engines… your content will not be used for AI model pre-training'. Sources are listed per answer, making it the most transparent of the major engines. Several studies find it leans on Reddit, LinkedIn and G2 for B2B queries.

OUR POSITION

⚠️ One unavoidable complication: on 4 August 2025 Cloudflare published allegations that Perplexity used stealth crawlers to bypass robots.txt and removed it from Verified Bots; Perplexity denied this. Recovery was still partial into early 2026, and **some Cloudflare zones still apply network-level reputation blocks — so allowing PerplexityBot in robots.txt may not be enough**.

01

What is actually documented

Perplexity states that PerplexityBot 'indexes pages similarly to other search engines… your content will not be used for AI model pre-training'. Sources are listed per answer, making it the most transparent of the major engines. Several studies find it leans on Reddit, LinkedIn and G2 for B2B queries.

On platform scale and share: Perplexity runs a Publisher Program revenue-share model, unlike the other major platforms.

Outside Google, no major platform has published anything about its citation algorithm. Every actionable playbook comes from correlation studies and practitioner testing — check the source behind any claim that a platform 'confirmed' something.

02

Citations are long-tail, not winner-takes-all

Measurements show even the most-cited domains rarely exceed 5% share on a single platform. The goal is reliable presence in the source list, not ownership of it.

The same question returns different source lists across time and region, so a single observation proves nothing. Sample repeatedly and read the distribution.

Data behind this page

under 5%

Ceiling on any single domain's citation share on one platform

SourceEvertune, 200M prompts,2026

40.1%

Reddit's share of citations in AI answers

SourceSemrush, 150,000+ citations across 5,000 keywords,2025-06

Sources

  1. [1]GEO: Generative Engine Optimization.Aggarwal et al., KDD 2024.2024

Updated 2026-08-06