← All posts

How to Block GPTBot Without Blocking Googlebot

Blocking AI crawlers is easy. Blocking them without hurting your search rankings is the part people get wrong.

The mistake usually looks like one of these: a blanket "block all bots" firewall rule that catches Googlebot in the net, a robots.txt that disallows everything, or a user-agent filter written in a hurry that matches more than intended. A week later, pages start dropping out of the index.

Here is how to separate the two cleanly.

Why they are different problems

Googlebot crawls your site to send you visitors. The relationship is an exchange: your content gets indexed, you get traffic. Blocking it costs you money.

GPTBot (and ClaudeBot, CCBot, Bytespider, and the rest) crawls your site to train or feed language models. There is no traffic coming back. Whether you allow it is a judgment call — some sites want the exposure inside AI answers, many don't want to donate their content — but it's your call, and it should be per-bot, not all-or-nothing.

Step 1: robots.txt — the polite request

User-agent: GPTBot
Disallow: /

User-agent: Googlebot
Allow: /

OpenAI says GPTBot respects robots.txt, and in practice it largely does. So do Googlebot, Bingbot, and most named AI crawlers. This is your first layer because it's free and standards-compliant.

But robots.txt is a request, not a wall. Bytespider is notorious for ignoring it, and scrapers that spoof GPTBot's user-agent never read it at all. Which brings us to layer two.

Step 2: verify, don't trust the user-agent string

Anyone can send User-Agent: GPTBot. Real verification means checking what the network says, not what the request claims:

  • Googlebot publishes its IP ranges and supports reverse-DNS verification (*.googlebot.com). A "Googlebot" request from a residential IP is a fake — block it with confidence.
  • OpenAI, Anthropic, and Common Crawl publish IP ranges for their crawlers too.
  • On Cloudflare-fronted infrastructure, the verified bot flag does this work for you at the edge: it confirms a crawler is who it claims to be before your rules even run.

The rule that follows: verified Googlebot passes, verified GPTBot gets your policy (allow or block), and unverified traffic claiming to be either is treated as a bad bot — because it is one.

Step 3: decide what "block" means

A 403 tells a scraper to rotate IPs and try again. Two stronger options:

  • Tarpit — serve the request, but slowly, drip-feeding bytes. The scraper's connection pool fills up with your molasses instead of moving to the next page. Cost lands on them, not you.
  • Maze / decoy content — serve generated filler pages so the crawl "succeeds" while collecting nothing real.

For verified, polite AI crawlers, a clean disallow is enough — they honor it. The tarpit is for the impolite ones.

Want the file written for you? The robots.txt generator builds it from per-bot checkboxes, and the crawler directory has the facts on each one.

The checklist

  1. robots.txt: disallow the AI crawlers you don't want, by name.
  2. Verify crawler identity by network, never by user-agent string alone.
  3. Per-bot policy: Googlebot always passes; each AI crawler gets an explicit allow/block.
  4. Unverified impostors → tarpit, not 403.
  5. Watch the live traffic for a week — you'll be surprised who's crawling.

TrafficDATA does all five out of the box: every visitor is classified (human, good bot, AI crawler, bad bot) with per-bot controls, verified-bot checks at the edge, and a tarpit engine for the ones that don't take no for an answer. The free tier covers 100K page views — enough to see exactly who's eating your bandwidth.