GLOSSARY

Training Crawler

IN ONE SENTENCE

A training crawler fetches content for model training, which has no direct bearing on whether you can be cited.

A training crawler fetches content for model training, which has no direct bearing on whether you can be cited.

01

In detail

Identifiers include GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Bytespider and CCBot. Whether to allow them is a brand policy choice. ⚠️ The worst configuration is blocking the retrieval crawlers while allowing the training ones: you contribute corpus and switch off exposure.

Updated 2026-08-06