anthropic-ai

Anthropic · Training claims robots compliance checked — limited vendor docs

What it does

anthropic-ai is Anthropic's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Legacy robots.txt token no longer listed in Anthropic's current bot roster (ClaudeBot superseded it). Kept because it still appears in older logs and block lists; near-zero live traffic expected.

Full UA: not captured in our July 2026 roster review — match log hits on the token anthropic-ai (case-insensitive substring) and verify by IP before acting.

How to verify a hit is really anthropic-ai

Verification method (July 2026): none documented (legacy token)

No published verification path. Anthropic publishes no IP list, rDNS pattern, or crawler-source documentation for anthropic-ai as of our July 2026 review — a UA hit cannot be positively confirmed as Anthropic infrastructure. Treat every anthropic-ai log line as unverified: UA strings are freely spoofed, and for this bot there is nothing to check them against. General triage still applies — see verify by IP.

Allow or block in robots.txt

Match the exact token anthropic-ai in robots.txt:

# Block anthropic-ai site-wide
User-agent: anthropic-ai
Disallow: /
# Explicitly allow anthropic-ai
User-agent: anthropic-ai
Allow: /
Compliance is claimed, not independently confirmed. Anthropic states this bot respects robots.txt; treat the directive as the first line of defense and watch your logs after deploying it.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether anthropic-ai may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.

Vendor documentation: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler

The Anthropic family

Anthropic operates 4 distinct tokens in this roster, and they do not block each other: a robots.txt line for anthropic-ai does nothing to its siblings — each needs its own User-agent: entry (the robots.txt guide has combined patterns).

TokenRolerobots.txt
ClaudeBotTraininghonors robots.txt
Claude-SearchBotAI searchhonors robots.txt
Claude-UserUser actionclaims robots compliance

Frequently asked questions

What is anthropic-ai?

anthropic-ai is Anthropic's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does anthropic-ai respect robots.txt?

Anthropic claims anthropic-ai respects robots.txt, but compliance is not independently confirmed.

How do I block anthropic-ai?

Add 'User-agent: anthropic-ai' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

How do I verify anthropic-ai traffic by IP?

You can't with confidence — Anthropic publishes no IP list or rDNS pattern for anthropic-ai as of our July 2026 review, so UA-only hits stay unverified.

Related bots

Same vendor, then other training crawlers — the full directory profiles all 34.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (25/34 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.