GLOSSARY
Training Crawler
IN ONE SENTENCE
A training crawler fetches content for model training, which has no direct bearing on whether you can be cited.
A training crawler fetches content for model training, which has no direct bearing on whether you can be cited.
01
In detail
Identifiers include GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Bytespider and CCBot. Whether to allow them is a brand policy choice. ⚠️ The worst configuration is blocking the retrieval crawlers while allowing the training ones: you contribute corpus and switch off exposure.
Updated 2026-08-06