Amazonbot

Amazon · Training honors robots.txt vendor-doc verified July 2026

What it does

Amazonbot is Amazon's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

'Used to improve our products and services' incl. Alexa — general-purpose crawl, closest bucket is training. Honors robots at each host level; caches robots.txt up to 30 days.

Full UA: not captured in our July 2026 roster review — match log hits on the token Amazonbot (case-insensitive substring) and verify by IP before acting.

How to verify a hit is really Amazonbot

Verification method (July 2026): Published IP list page: https://developer.amazon.com/amazonbot/ip-addresses (HTML page, not raw JSON)

The vendor publishes crawler IPs on a documentation page (not a machine-readable JSON endpoint), so verification is a manual cross-check: open the page, compare the hit's source IP against the listed ranges.

User-agent strings are freely spoofed, so a UA match alone proves nothing. Full step-by-step method: verify by IP.

Allow or block in robots.txt

Match the exact token Amazonbot in robots.txt:

# Block Amazonbot site-wide
User-agent: Amazonbot
Disallow: /
# Explicitly allow Amazonbot
User-agent: Amazonbot
Allow: /
Compliant bots stop fast. Well-behaved crawlers stop within roughly one crawl cycle of a robots.txt Disallow (we observed GPTBot's hammering on a production site we operate stop within one cycle). Give a fresh directive about a day before concluding it's being ignored.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether Amazonbot may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.

Vendor documentation: https://developer.amazon.com/amazonbot

The Amazon family

Amazon operates 3 distinct tokens in this roster, and they do not block each other: a robots.txt line for Amazonbot does nothing to its siblings — each needs its own User-agent: entry (the robots.txt guide has combined patterns).

TokenRolerobots.txt
Amzn-SearchBotAI searchhonors robots.txt
Amzn-UserUser actionclaims robots compliance

Frequently asked questions

What is Amazonbot?

Amazonbot is Amazon's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does Amazonbot respect robots.txt?

Yes — Amazon documents robots.txt compliance for Amazonbot (vendor docs reviewed July 2026).

How do I block Amazonbot?

Add 'User-agent: Amazonbot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

How do I verify Amazonbot traffic by IP?

Amazon publishes Amazonbot's IP ranges on a documentation page (not machine-readable JSON); cross-check the hit's source IP against the listed ranges manually.

Related bots

Same vendor, then other training crawlers — the full directory profiles all 34.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (25/34 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.