Crawler data methodology & sources

Where the bot facts come from. Every bot fact in this section — tokens, purposes, robots.txt stances, IP lists — comes from one dataset compiled July 2026 from vendor crawler documentation, published IP-range JSONs (fetched live 2026-07-19), and field observations from production access logs and block lists on sites we operate. 21 of 28 bots are vendor-doc verified: confirmed against the vendor's own current crawler documentation this review. The remaining 7 are checked: the vendor publishes no usable docs (Bytespider, cohere-ai), the docs deep-link 404'd (Diffbot), or the docs page refused our fetch — Meta's web-crawlers page returned HTTP 400 to non-browser requests this pass, so both Meta bots carry "checked" and their robots-compliance claims are labeled as claims. Every bot page shows which tier it's in.

How the analyzer works. Entirely in your browser: JavaScript splits pasted lines, extracts the user-agent field (last quoted field of common/combined log format that isn't the request line or a referrer URL, or the whole line for plain UA pastes; a line whose UA field is "-" has no user agent and is never matched), strips any "+http…" vendor URL from the UA, and case-insensitively substring-matches against the 26 log-matchable tokens. Nothing is transmitted — no upload endpoint exists for the analyzer. Two roster entries (Google-Extended, Applebot-Extended) are excluded from matching because they have no user agent; lines carrying them are flagged as spoofed.

Limits. UA matching is triage, not proof — spoofers borrow bot names, and some bots rotate UAs. Verify by IP before acting on a hit. The roster reflects July 2026; AI-crawler tokens churn fast (new hyphenated tokens don't match old patterns), and we re-verify on that cadence.

Vendor documentation reviewed (July 2026):

Post-launch: automated crawler monitoring. A watch service — weekly reports of which AI bots hit your site, IP-verified, with roster updates as vendors ship new tokens — is planned but not yet available. Get notified when it ships →

Corrections. Spotted a wrong token, a stale robots stance, a bot we're missing? Email us — fixes ship within a day.

Disclosure. The robots.txt guide carries hosting/CDN links that may earn us a commission. Commissions never alter bot facts — those are sourced above so you can check us. Independent and solo-operated.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (21/28 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.