AI2Bot

Allen Institute for AI · Training honors robots.txt vendor-doc verified July 2026

What it does

AI2Bot is Allen Institute for AI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Crawls for open language-model research (OLMo/Dolma). UA: Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler). A Dolma-specific variant 'AI2Bot-Dolma' also exists — prefix-match AI2Bot to catch both.

How to verify a hit is really AI2Bot

Verification method (July 2026): none documented (no published IP list)

User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to Allen Institute for AI — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.

Allow or block in robots.txt

Match the exact token AI2Bot in robots.txt:

# Block AI2Bot site-wide
User-agent: AI2Bot
Disallow: /
# Explicitly allow AI2Bot
User-agent: AI2Bot
Allow: /
Compliant bots stop fast. Well-behaved crawlers stop within roughly one crawl cycle of a robots.txt Disallow (we observed GPTBot's hammering on a production site we operate stop within one cycle). Give a fresh directive about a day before concluding it's being ignored.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.

Vendor documentation: https://allenai.org/crawler

Frequently asked questions

What is AI2Bot?

AI2Bot is Allen Institute for AI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does AI2Bot respect robots.txt?

Yes — Allen Institute for AI documents robots.txt compliance for AI2Bot (vendor docs reviewed July 2026).

How do I block AI2Bot?

Add 'User-agent: AI2Bot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

Other bots

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (21/28 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.