CCBot
Builds the open Common Crawl corpus, which many AI labs use as training data downstream. One robots.txt line here affects a wide swath of future training sets.
| robots.txt token | CCBot |
| Respects robots.txt | Yes — long, consistent compliance record |
| Verification | Operated by the Common Crawl non-profit; documented crawler |
| Our recommendation | Your call — The highest-leverage single decision: blocking CCBot removes you from many datasets at once — in either direction, make it deliberately. |
Block CCBot with robots.txt
User-agent: CCBot
Disallow: /robots.txt is a request, not a wall — requests claiming to be CCBot from unverified networks should be treated as bad bots. How to verify crawlers by network →
See every CCBot request hitting your site — live, with per-bot allow / block / tarpit controls. Try TrafficDATA free (100K page views)