What it does
CCBot is Common Crawl's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Non-profit; corpus is the seed data for many LLM training sets — blocking CCBot is the broadest single training opt-out. Common Crawl itself warns spoofers impersonate CCBot: verify by IP/rDNS. UA: CCBot/2.0 (https://commoncrawl.org/faq/)
Full UA: not captured in our July 2026 roster review — match log hits on the token CCBot (case-insensitive substring) and verify by IP before acting.
How to verify a hit is really CCBot
Verification method (July 2026): Published IP list: https://index.commoncrawl.org/ccbot.json (fetched live 2026-07-19) + rDNS *.crawl.commoncrawl.org (IPv4)
User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to Common Crawl — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.
Allow or block in robots.txt
Match the exact token CCBot in robots.txt:
# Block CCBot site-wide
User-agent: CCBot
Disallow: /
# Explicitly allow CCBot
User-agent: CCBot
Allow: /
For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.
Vendor documentation: https://commoncrawl.org/ccbot
Frequently asked questions
What is CCBot?
CCBot is Common Crawl's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Does CCBot respect robots.txt?
Yes — Common Crawl documents robots.txt compliance for CCBot (vendor docs reviewed July 2026).
How do I block CCBot?
Add 'User-agent: CCBot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.