AI2Bot

Allen Institute for AI · Training honors robots.txt vendor-doc verified July 2026

What it does

AI2Bot is Allen Institute for AI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Crawls for open language-model research (OLMo/Dolma). UA: Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler). The 'AI2Bot-Dolma' variant no longer appears on the vendor page (checked 2026-08-22) but exists in older logs — prefix-match AI2Bot to catch both.

How to verify a hit is really AI2Bot

Verification method (July 2026): none documented (no published IP list)

No published verification path. Allen Institute for AI publishes no IP list, rDNS pattern, or crawler-source documentation for AI2Bot as of our July 2026 review — a UA hit cannot be positively confirmed as Allen Institute for AI infrastructure. Treat every AI2Bot log line as unverified: UA strings are freely spoofed, and for this bot there is nothing to check them against. General triage still applies — see verify by IP.

Allow or block in robots.txt

Match the exact token AI2Bot in robots.txt:

# Block AI2Bot site-wide
User-agent: AI2Bot
Disallow: /
# Explicitly allow AI2Bot
User-agent: AI2Bot
Allow: /
Compliant bots stop fast. Well-behaved crawlers stop within roughly one crawl cycle of a robots.txt Disallow (we observed GPTBot's hammering on a production site we operate stop within one cycle). Give a fresh directive about a day before concluding it's being ignored.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether AI2Bot may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.

Vendor documentation: https://allenai.org/crawler

The Allen Institute for AI family

AI2Bot is the only Allen Institute for AI token in our 34-bot roster — one robots.txt line covers everything Allen Institute for AI operates here.

Frequently asked questions

What is AI2Bot?

AI2Bot is Allen Institute for AI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does AI2Bot respect robots.txt?

Yes — Allen Institute for AI documents robots.txt compliance for AI2Bot (vendor docs reviewed July 2026).

How do I block AI2Bot?

Add 'User-agent: AI2Bot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

How do I verify AI2Bot traffic by IP?

You can't with confidence — Allen Institute for AI publishes no IP list or rDNS pattern for AI2Bot as of our July 2026 review, so UA-only hits stay unverified.

Related bots

Same vendor, then other training crawlers — the full directory profiles all 34.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (25/34 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.