Crawler Directory

Who's Crawling the Web

Every notable crawler — what it collects, whether it honors robots.txt, how to verify it, and whether to let it in. Generate a robots.txt from any of them in one click.

AI Training

OpenAIYour call

GPTBot

Crawls public web pages to collect training data for OpenAI's GPT models. Allowing it means your content can inform future models; there is no traffic in return.

AnthropicYour call

ClaudeBot

Collects public web content for training and improving Anthropic's Claude models.

Common CrawlYour call

CCBot

Builds the open Common Crawl corpus, which many AI labs use as training data downstream. One robots.txt line here affects a wide swath of future training sets.

GoogleYour call

Google-Extended

Not a separate crawler — a robots.txt token that controls whether content Googlebot already crawls may also be used for Gemini training. Disallowing it does NOT affect your search indexing.

AppleYour call

Applebot-Extended

Like Google-Extended: a robots.txt token controlling whether Applebot's crawl may feed Apple's AI training. Disallowing it keeps you in Siri/Spotlight search.

ByteDanceBlock

Bytespider

ByteDance's data crawler, widely associated with AI training collection. Notorious for aggressive crawl volumes.

MetaYour call

Meta-ExternalAgent

Meta's crawler for collecting web content used in AI model training and product improvement.

AmazonYour call

Amazonbot

Crawls for Amazon services including Alexa answers and AI features.

Allen Institute for AIYour call

AI2Bot

Collects web content for the Allen Institute's open AI research models.

Webz.ioBlock

Omgilibot / Webzio

A commercial data broker crawl whose feeds are resold, including into AI training pipelines.

HiveYour call

ImagesiftBot

Collects images from the web for Hive's visual-AI products.

HuaweiYour call

PetalBot

Crawls for Huawei's Petal Search and related AI features.

DiffbotYour call

Diffbot

Extracts structured data from web pages into a commercial knowledge graph, resold via API to customers including AI products.

GoogleYour call

GoogleOther

Google's generic crawler for internal research, development, and one-off product fetches. Explicitly NOT used for Search ranking — blocking it does not affect your indexing.