Who's Crawling the Web
Every notable crawler — what it collects, whether it honors robots.txt, how to verify it, and whether to let it in. Generate a robots.txt from any of them in one click.
AI Training
GPTBot
Crawls public web pages to collect training data for OpenAI's GPT models. Allowing it means your content can inform future models; there is no traffic in return.
ClaudeBot
Collects public web content for training and improving Anthropic's Claude models.
CCBot
Builds the open Common Crawl corpus, which many AI labs use as training data downstream. One robots.txt line here affects a wide swath of future training sets.
Google-Extended
Not a separate crawler — a robots.txt token that controls whether content Googlebot already crawls may also be used for Gemini training. Disallowing it does NOT affect your search indexing.
Applebot-Extended
Like Google-Extended: a robots.txt token controlling whether Applebot's crawl may feed Apple's AI training. Disallowing it keeps you in Siri/Spotlight search.
Bytespider
ByteDance's data crawler, widely associated with AI training collection. Notorious for aggressive crawl volumes.
Meta-ExternalAgent
Meta's crawler for collecting web content used in AI model training and product improvement.
Amazonbot
Crawls for Amazon services including Alexa answers and AI features.
AI2Bot
Collects web content for the Allen Institute's open AI research models.
Omgilibot / Webzio
A commercial data broker crawl whose feeds are resold, including into AI training pipelines.
ImagesiftBot
Collects images from the web for Hive's visual-AI products.
PetalBot
Crawls for Huawei's Petal Search and related AI features.
Diffbot
Extracts structured data from web pages into a commercial knowledge graph, resold via API to customers including AI products.
GoogleOther
Google's generic crawler for internal research, development, and one-off product fetches. Explicitly NOT used for Search ranking — blocking it does not affect your indexing.
AI Assistant
ChatGPT-User
Fetches a page live when a ChatGPT user asks about it — not bulk crawling. Blocking it stops ChatGPT users from reading your pages on demand.
Claude-User
Fetches pages on demand when a Claude user asks about them.
Meta-ExternalFetcher
Fetches individual links on direct user request inside Meta products.
DuckAssistBot
Retrieves pages to generate DuckAssist AI answers, with source citations.
Perplexity-User
Fetches a page live when a Perplexity user asks about it. Distinct from PerplexityBot (the index crawler) — this one only arrives on user request.
AI Search
Search Engine
Googlebot
Indexes your site for Google Search. The crawler your traffic depends on.
Bingbot
Indexes for Bing — which also feeds several AI assistants' search layers.
Applebot
Crawls for Siri and Spotlight suggestions. AI-training use is controlled separately via Applebot-Extended.
DuckDuckBot
Crawls for DuckDuckGo search results.
YandexBot
Indexes for Yandex, the dominant search engine for Russian-language audiences.
Baiduspider
Indexes for Baidu, the leading search engine in China.
Sogou Spider
Indexes for Sogou, the Chinese search engine integrated into Tencent's WeChat ecosystem.
YisouSpider
Crawls for Shenma, Alibaba's mobile search engine in the UC Browser ecosystem.
360Spider
Indexes for so.com, Qihoo 360's search engine.
SEO Tool
AhrefsBot
Builds Ahrefs' backlink index. Heavy crawler; useful to you only if you (or your competitors' analysts) use Ahrefs.
SemrushBot
Crawls for Semrush's SEO index.
MJ12bot
Builds Majestic's backlink index. Unusually, it's a distributed crawler running on volunteer machines — so genuine MJ12bot requests routinely arrive from residential IPs, which looks like spoofing but isn't.
DotBot
Crawls for Moz's link index (the data behind Domain Authority).
Link Preview
facebookexternalhit
Fetches a page's Open Graph tags to render the preview card when someone shares your URL on Facebook, Messenger, or Instagram.
Twitterbot
Fetches Twitter Card metadata to render link previews when your URL is posted on X.
LinkedInBot
Fetches metadata for link previews when your URL is shared on LinkedIn.
Slackbot
Unfurls links posted in Slack channels into preview cards.
Discordbot
Fetches link previews for URLs posted in Discord servers and DMs.
TelegramBot
Fetches link previews (and Instant View versions) for URLs shared in Telegram chats.
Fetches link previews when a URL is shared in a WhatsApp chat.
Scanner
CensysInspect
Scans the public internet to map exposed services for Censys's security search engine. It arrives at your server because your IP exists, not because anything linked to you.
Shadowserver
A security non-profit that scans the internet for vulnerable and exposed services, then notifies network owners for free.
Expanse (Cortex Xpanse)
Scans the internet to map organizations' attack surfaces for Palo Alto's commercial Cortex Xpanse product. Its UA states its identity and contact in full.
InternetMeasurement
Driftnet's internet-measurement scans, documented with a contact page in the UA.
Automation Tool
cURL
The standard command-line HTTP client. In your logs it means a script or a person fetched the page outside a browser — uptime checks, scrapers, and exploit probes all look the same here.
python-requests
Python's most popular HTTP library, and the default UA of half the quick scraping scripts ever written.
Go-http-client
The default UA of Go's standard HTTP client — common in monitoring agents, custom crawlers, and bulk fetchers.
Scrapy
A purpose-built web-scraping framework. Unlike generic HTTP clients, a Scrapy UA means someone is systematically extracting your content.
HeadlessChrome
Real Chrome running without a window — the engine under Puppeteer and Playwright. It executes JavaScript and looks like a browser because it is one; the UA token is the honest default that automation forgot to remove.