What it does
Bytespider is ByteDance's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
TikTok/Doubao LLM training crawler. No official documentation; widely reported to ignore robots.txt and rotate UAs. Blocking requires WAF rules. Present in our production block lists on both operator properties.
Full UA: not captured in our July 2026 roster review — match log hits on the token Bytespider (case-insensitive substring) and verify by IP before acting.
How to verify a hit is really Bytespider
Verification method (July 2026): none documented (no official docs, no published IP ranges)
User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to ByteDance — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.
Allow or block in robots.txt
Match the exact token Bytespider in robots.txt:
# Block Bytespider site-wide
User-agent: Bytespider
Disallow: /
# Explicitly allow Bytespider
User-agent: Bytespider
Allow: /
For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.
No official vendor documentation is published for this bot.
Frequently asked questions
What is Bytespider?
Bytespider is ByteDance's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Does Bytespider respect robots.txt?
No — Bytespider does not treat robots.txt as binding. Blocking it requires WAF or IP-level rules.
How do I block Bytespider?
robots.txt cannot stop Bytespider; block it with WAF or IP-level firewall rules instead.