Diffbot

Diffbot · Training ignores robots.txt checked — limited vendor docs

What it does

Diffbot is Diffbot's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Commercial extraction/Knowledge-Graph crawler; fetches on behalf of paying customers and has stated it does not treat robots.txt as binding for customer-directed fetches. Docs deep-link 404'd this pass — treat as checked.

Full UA: not captured in our July 2026 roster review — match log hits on the token Diffbot (case-insensitive substring) and verify by IP before acting.

How to verify a hit is really Diffbot

Verification method (July 2026): none documented

No published verification path. Diffbot publishes no IP list, rDNS pattern, or crawler-source documentation for Diffbot as of our July 2026 review — a UA hit cannot be positively confirmed as Diffbot infrastructure. Treat every Diffbot log line as unverified: UA strings are freely spoofed, and for this bot there is nothing to check them against. General triage still applies — see verify by IP.

Allow or block in robots.txt

Match the exact token Diffbot in robots.txt:

# Block Diffbot site-wide
User-agent: Diffbot
Disallow: /
# Explicitly allow Diffbot
User-agent: Diffbot
Allow: /
robots.txt will not stop Diffbot. Per the vendor's own documentation or observed behavior, this bot does not treat robots.txt as binding. If you need it stopped, use WAF or IP-level rules — see verifying and blocking by IP.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether Diffbot may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.

Vendor documentation: https://www.diffbot.com/docs/

The Diffbot family

Diffbot is the only Diffbot token in our 34-bot roster — one robots.txt line covers everything Diffbot operates here.

Frequently asked questions

What is Diffbot?

Diffbot is Diffbot's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does Diffbot respect robots.txt?

No — Diffbot does not treat robots.txt as binding. Blocking it requires WAF or IP-level rules.

How do I block Diffbot?

robots.txt cannot stop Diffbot; block it with WAF or IP-level firewall rules instead.

How do I verify Diffbot traffic by IP?

You can't with confidence — Diffbot publishes no IP list or rDNS pattern for Diffbot as of our July 2026 review, so UA-only hits stay unverified.

Related bots

Same vendor, then other training crawlers — the full directory profiles all 34.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (25/34 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.