← All posts

What Is a Bot Tarpit? (And Why a 403 Doesn't Stop Scrapers)

When you block a scraper with a 403, you're giving it useful information: this IP is burned, rotate and continue. Modern scraping stacks treat blocks as routine — they run through residential proxy pools and a 403 costs them milliseconds.

A tarpit takes the opposite approach: never say no — say yes, very slowly.

How a tarpit works

Instead of rejecting the request, the server accepts it and responds at a crawl: a few bytes, a pause, a few more bytes. The connection stays technically alive, so the client doesn't error out — it waits.

One tarpitted request is nothing. But scrapers are built for throughput: they hold connection pools, and every connection stuck in your tarpit is one that can't fetch the next page. Pool fills up, throughput collapses, and the scrape that should have taken minutes takes hours — or gets aborted by its own timeout logic.

The asymmetry is the point. Serving bytes slowly costs you almost nothing (on edge infrastructure, effectively nothing — the connection holds at the PoP, not your origin). The scraper pays in time, connection slots, and proxy bandwidth — the things it actually budgets for.

Tarpit vs. the alternatives

ResponseScraper's experienceYour cost
403 block"Rotate IP, retry" — seconds lostNone, but no deterrence
CAPTCHASolved by services for ~$1/1000Annoys your human users too
Rate limitBacks off, returns slowerScrape still completes
TarpitConnections hang, throughput dies~Nothing at the edge

The 403 also has a subtler failure: it tells the scraper which requests got detected, which is free feedback for tuning evasion. A tarpit gives no clean signal — was the site slow? overloaded? — so there's nothing to tune against.

When not to tarpit

Don't tarpit anything you might want back. A misclassified human (false positive) stuck in a tarpit has a terrible experience with no explanation — which is why tarpitting should sit behind classification you trust, applied only to traffic that's verified bad: exploit probes, robots.txt-ignoring crawlers, impostor user-agents. Polite bots that honor a disallow never need it.

And never tarpit verified search engine crawlers, obviously — that's self-inflicted SEO damage in slow motion.

The short version

A 403 is a door slam a scraper barely notices. A tarpit is quicksand: invisible until they're in it, expensive to leave, nothing to learn from. It flips the economics — scraping you stops being free.

TrafficDATA ships a tarpit engine as a per-bot policy: classify first (human / good bot / AI crawler / bad bot), then choose monitor, block, honeypot, or tarpit per domain. You can watch tarpitted connections drain scraper patience in the live feed, in real time.