If you haven't looked at your server logs lately, here's the short version: a meaningful share of your traffic is now AI crawlers, and most site owners have never made a deliberate decision about any of them.
Here are the ones that matter, what they're for, and how they behave.
The named crawlers
GPTBot (OpenAI) — collects training data for GPT models. Honors robots.txt, publishes its IP ranges. The one most people ask about first.
OAI-SearchBot / ChatGPT-User (OpenAI) — different jobs: search indexing for ChatGPT's browsing, and live fetches when a user asks ChatGPT to read a page. Blocking GPTBot does not block these — they have their own user-agents, which is exactly why per-bot policy beats one big switch.
ClaudeBot (Anthropic) — training and retrieval data for Claude. Honors robots.txt, publishes ranges.
Google-Extended (Google) — not a separate crawler: it's a robots.txt token that controls whether regular Googlebot's crawl may also feed Gemini training. Disallowing Google-Extended does not affect your search indexing — that's the entire point of it existing.
CCBot (Common Crawl) — builds the public web corpus that many models train on downstream. Blocking CCBot cuts your content out of a wide swath of future training sets in one move. Honors robots.txt.
PerplexityBot (Perplexity) — feeds an AI answer engine that does cite sources, so there's a real (if modest) traffic-back argument here.
Bytespider (ByteDance) — the misbehaver of the list. Widely reported to ignore robots.txt and crawl aggressively. Treat it as a bad bot: this is tarpit territory, not polite-disallow territory.
The unnamed ones
The named crawlers are the visible part. Below them is a layer of scrapers that spoof browser user-agents or impersonate the named bots to ride on their reputation. No robots.txt line helps there — you need network-level verification (published IP ranges, reverse DNS, edge verified-bot flags) and behavioral signals: request pacing, TLS fingerprints, whether "Chrome" actually executes JavaScript.
A simple rule of thumb: a crawler that tells you its name usually takes no for an answer. The ones that lie about their name are the ones worth defending against.
A sane default policy
- Decide the easy ones: Googlebot/Bingbot always allowed (that's your search traffic). Bytespider blocked or tarpitted (it doesn't ask).
- Make the judgment calls per-bot: GPTBot, ClaudeBot, CCBot — allow if you want presence inside AI answers and training sets, disallow if you don't want to donate content. Either answer is defensible; not choosing is the only wrong move.
- Verify identity by network, not user-agent — impostors inherit the bad bot policy automatically.
- Re-check quarterly. New crawlers appear constantly; your logs are the only honest census.
Want the full census with a per-bot verdict? The crawler directory covers 25+ of them, and the robots.txt generator turns your decisions into a paste-ready file.
That census is the part most people skip, because grepping logs is miserable. TrafficDATA does it continuously — every AI crawler identified and badged in a real-time feed, with per-bot allow/block/tarpit controls per domain. The free 100K-view tier is plenty to find out who's already there.