ClaudeBot

Anthropic · Training honors robots.txt vendor-doc verified July 2026

What it does

ClaudeBot is Anthropic's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Anthropic documents robots.txt + Crawl-delay support and no-CAPTCHA-circumvention. Blocking by IP can break the opt-out (impedes robots.txt reads) — prefer robots directives.

Full UA: not captured in our July 2026 roster review — match log hits on the token ClaudeBot (case-insensitive substring) and verify by IP before acting.

How to verify a hit is really ClaudeBot

Verification method (July 2026): Published IP list: https://claude.com/crawling/bots.json (fetched live 2026-07-19)

User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to Anthropic — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.

Allow or block in robots.txt

Match the exact token ClaudeBot in robots.txt:

# Block ClaudeBot site-wide
User-agent: ClaudeBot
Disallow: /
# Explicitly allow ClaudeBot
User-agent: ClaudeBot
Allow: /
Compliant bots stop fast. Well-behaved crawlers stop within roughly one crawl cycle of a robots.txt Disallow (we observed GPTBot's hammering on a production site we operate stop within one cycle). Give a fresh directive about a day before concluding it's being ignored.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.

Vendor documentation: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler

Frequently asked questions

What is ClaudeBot?

ClaudeBot is Anthropic's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does ClaudeBot respect robots.txt?

Yes — Anthropic documents robots.txt compliance for ClaudeBot (vendor docs reviewed July 2026).

How do I block ClaudeBot?

Add 'User-agent: ClaudeBot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

Other bots

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (21/28 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.