PLATFORMS

How DeepSeek Decides What to Cite

IN ONE SENTENCE

Getting cited in DeepSeek requires content to be inside its retrievable set and, for a given question, more worth citing than the alternatives.

Being named in DeepSeek requires two things: the content is inside its retrievable set, and for a given question it is more worth citing than the alternatives. Practitioner studies of citation mix generally find DeepSeek leaning on news media plus Chinese UGC such as Zhihu. It publishes no dedicated retrieval crawler identifier and does not display sources per answer, so visibility can only be measured by sampling answers against a fixed question set.

OUR POSITION

No published crawler identifier does not mean no crawling. Look for it in server logs rather than assuming absence because no matching user agent appears in robots.txt.

01

What is actually documented

Practitioner studies of citation mix generally find DeepSeek leaning on news media plus Chinese UGC such as Zhihu. It publishes no dedicated retrieval crawler identifier and does not display sources per answer, so visibility can only be measured by sampling answers against a fixed question set.

On platform scale and share: DeepSeek ranks near the top of China's AI-native apps by MAU, making it as essential as Doubao for the domestic market.

Outside Google, no major platform has published anything about its citation algorithm. Every actionable playbook comes from correlation studies and practitioner testing — check the source behind any claim that a platform 'confirmed' something.

02

Citations are long-tail, not winner-takes-all

Measurements show even the most-cited domains rarely exceed 5% share on a single platform. The goal is reliable presence in the source list, not ownership of it.

The same question returns different source lists across time and region, so a single observation proves nothing. Sample repeatedly and read the distribution.

Data behind this page

under 5%

Ceiling on any single domain's citation share on one platform

SourceEvertune, 200M prompts,2026

40.1%

Reddit's share of citations in AI answers

SourceSemrush, 150,000+ citations across 5,000 keywords,2025-06

Sources

  1. [1]GEO: Generative Engine Optimization.Aggarwal et al., KDD 2024.2024

Updated 2026-08-06