What it does
CCBot is Common Crawl's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Non-profit; corpus is the seed data for many LLM training sets — blocking CCBot is the broadest single training opt-out. Common Crawl itself warns spoofers impersonate CCBot: verify by IP/rDNS. UA: CCBot/2.0 (https://commoncrawl.org/faq/)
Full UA: not captured in our July 2026 roster review — match log hits on the token CCBot (case-insensitive substring) and verify by IP before acting.
How to verify a hit is really CCBot
Verification method (July 2026): Published IP list: https://index.commoncrawl.org/ccbot.json (fetched live 2026-08-22) + rDNS *.crawl.commoncrawl.org (IPv4)
Strongest available check: the vendor publishes the crawler's egress IPs. Fetch the list, then confirm the hit's source IP falls inside it — a UA match from any other IP is a spoofer.
User-agent strings are freely spoofed, so a UA match alone proves nothing. Full step-by-step method: verify by IP.
Allow or block in robots.txt
Match the exact token CCBot in robots.txt:
# Block CCBot site-wide
User-agent: CCBot
Disallow: /
# Explicitly allow CCBot
User-agent: CCBot
Allow: /
For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether CCBot may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.
Vendor documentation: https://commoncrawl.org/ccbot
The Common Crawl family
CCBot is the only Common Crawl token in our 34-bot roster — one robots.txt line covers everything Common Crawl operates here.
Frequently asked questions
What is CCBot?
CCBot is Common Crawl's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Does CCBot respect robots.txt?
Yes — Common Crawl documents robots.txt compliance for CCBot (vendor docs reviewed July 2026).
How do I block CCBot?
Add 'User-agent: CCBot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.
How do I verify CCBot traffic by IP?
Fetch Common Crawl's published IP list for CCBot and confirm the hit's source IP falls inside it — a CCBot user-agent from any other IP is a spoofer.
Related bots
Same vendor, then other training crawlers — the full directory profiles all 34.