Bytespider

ByteDance · Training ignores robots.txt checked — limited vendor docs

What it does

Bytespider is ByteDance's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

TikTok/Doubao LLM training crawler. No official documentation; widely reported to ignore robots.txt and rotate UAs. Blocking requires WAF rules. Present in our production block lists on both operator properties.

Full UA: not captured in our July 2026 roster review — match log hits on the token Bytespider (case-insensitive substring) and verify by IP before acting.

How to verify a hit is really Bytespider

Verification method (July 2026): none documented (no official docs, no published IP ranges)

User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to ByteDance — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.

Allow or block in robots.txt

Match the exact token Bytespider in robots.txt:

# Block Bytespider site-wide
User-agent: Bytespider
Disallow: /
# Explicitly allow Bytespider
User-agent: Bytespider
Allow: /
robots.txt will not stop Bytespider. Per the vendor's own documentation or observed behavior, this bot does not treat robots.txt as binding. If you need it stopped, use WAF or IP-level rules — see verifying and blocking by IP.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.

No official vendor documentation is published for this bot.

Frequently asked questions

What is Bytespider?

Bytespider is ByteDance's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does Bytespider respect robots.txt?

No — Bytespider does not treat robots.txt as binding. Blocking it requires WAF or IP-level rules.

How do I block Bytespider?

robots.txt cannot stop Bytespider; block it with WAF or IP-level firewall rules instead.

Other bots

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (21/28 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.