GoogleOther

Google · Training honors robots.txt vendor-doc verified July 2026

What it does

GoogleOther is Google's generic crawler for internal research and development — Google states it is NOT used for Search indexing, and Google's AI-training control is the separate Google-Extended robots token. We file it under 'training' as the closest bucket, but blocking Google-Extended (not GoogleOther) is how you opt out of Gemini training.

Generic Google crawler for internal R&D / one-off fetches; not a ranking crawler. Categorized under training as the closest bucket — blocking it does not affect Search.

Full UA: not captured in our July 2026 roster review — match log hits on the token GoogleOther (case-insensitive substring) and verify by IP before acting.

How to verify a hit is really GoogleOther

Verification method (July 2026): Published IP list: https://developers.google.com/static/crawling/ipranges/common-crawlers.json (fetched live 2026-08-22); old search/apis/ipranges/googlebot.json path 301s there. Plus rDNS crawl-*.googlebot.com. GoogleOther-Image / GoogleOther-Video variants documented on the same page.

Strongest available check: the vendor publishes the crawler's egress IPs. Fetch the list, then confirm the hit's source IP falls inside it — a UA match from any other IP is a spoofer.

User-agent strings are freely spoofed, so a UA match alone proves nothing. Full step-by-step method: verify by IP.

Allow or block in robots.txt

Match the exact token GoogleOther in robots.txt:

# Block GoogleOther site-wide
User-agent: GoogleOther
Disallow: /
# Explicitly allow GoogleOther
User-agent: GoogleOther
Allow: /
Compliant bots stop fast. Well-behaved crawlers stop within roughly one crawl cycle of a robots.txt Disallow (we observed GPTBot's hammering on a production site we operate stop within one cycle). Give a fresh directive about a day before concluding it's being ignored.

For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide. robots.txt governs whether GoogleOther may fetch; llms.txt is the separate, curated map AI assistants read once allowed in.

Vendor documentation: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers

The Google family

Google operates 5 distinct tokens in this roster, and they do not block each other: a robots.txt line for GoogleOther does nothing to its siblings — each needs its own User-agent: entry (the robots.txt guide has combined patterns).

TokenRolerobots.txt
Google-ExtendedTraininghonors robots.txt
Google-CloudVertexBotUser actionhonors robots.txt
Google-AgentUser actionignores robots.txt
Google-GeminiNotebookUser actionignores robots.txt

Frequently asked questions

What is GoogleOther?

GoogleOther is Google's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.

Does GoogleOther respect robots.txt?

Yes — Google documents robots.txt compliance for GoogleOther (vendor docs reviewed July 2026).

How do I block GoogleOther?

Add 'User-agent: GoogleOther' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.

How do I verify GoogleOther traffic by IP?

Fetch Google's published IP list for GoogleOther and confirm the hit's source IP falls inside it — a GoogleOther user-agent from any other IP is a spoofer.

Related bots

Same vendor, then other training crawlers — the full directory profiles all 34.

AnswerFootprint crawler analytics — bot roster compiled July 2026 from vendor crawler docs and published IP lists (25/34 vendor-doc verified). The analyzer runs 100% client-side — your logs never leave this browser. User-agent strings can be spoofed; treat UA matches as a first pass and verify by IP before acting. Errors in the bot data: report them.