Where the bot facts come from. Every bot fact in this section — tokens, purposes, robots.txt stances, IP lists — comes from one dataset compiled July 2026 from vendor crawler documentation, published IP-range JSONs (fetched live 2026-07-19), and field observations from production access logs and block lists on sites we operate. 21 of 28 bots are vendor-doc verified: confirmed against the vendor's own current crawler documentation this review. The remaining 7 are checked: the vendor publishes no usable docs (Bytespider, cohere-ai), the docs deep-link 404'd (Diffbot), or the docs page refused our fetch — Meta's web-crawlers page returned HTTP 400 to non-browser requests this pass, so both Meta bots carry "checked" and their robots-compliance claims are labeled as claims. Every bot page shows which tier it's in.
How the analyzer works. Entirely in your browser: JavaScript splits pasted lines, extracts the user-agent field (last quoted field of common/combined log format that isn't the request line or a referrer URL, or the whole line for plain UA pastes; a line whose UA field is "-" has no user agent and is never matched), strips any "+http…" vendor URL from the UA, and case-insensitively substring-matches against the 26 log-matchable tokens. Nothing is transmitted — no upload endpoint exists for the analyzer. Two roster entries (Google-Extended, Applebot-Extended) are excluded from matching because they have no user agent; lines carrying them are flagged as spoofed.
Limits. UA matching is triage, not proof — spoofers borrow bot names, and some bots rotate UAs. Verify by IP before acting on a hit. The roster reflects July 2026; AI-crawler tokens churn fast (new hyphenated tokens don't match old patterns), and we re-verify on that cadence.
Vendor documentation reviewed (July 2026):
- https://platform.openai.com/docs/bots (fetched 2026-07-19)
- https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (fetched 2026-07-19)
- https://docs.perplexity.ai/guides/bots (fetched 2026-07-19)
- https://developer.amazon.com/amazonbot (fetched 2026-07-19)
- https://support.apple.com/en-us/119829 (fetched 2026-07-19)
- https://commoncrawl.org/ccbot (fetched 2026-07-19)
- https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/ (fetched 2026-07-19)
- https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (fetched 2026-07-19)
- https://allenai.org/crawler (fetched 2026-07-19)
- https://docs.mistral.ai/robots/ (fetched 2026-07-19)
Corrections. Spotted a wrong token, a stale robots stance, a bot we're missing? Email us — fixes ship within a day.
Disclosure. The robots.txt guide carries hosting/CDN links that may earn us a commission. Commissions never alter bot facts — those are sourced above so you can check us. Independent and solo-operated.