What it does
Google-Extended never appears in access logs. It has no HTTP user-agent string — it exists only as a robots.txt User-agent: token. The actual crawling happens under Google's regular crawler user agents; this token controls whether the content those crawlers already fetch may be used for AI training.
NEVER appears in access logs. Google-Extended has no HTTP user-agent string — crawling happens under regular Googlebot UAs; the token only works as a robots.txt User-agent line controlling Gemini training / AI grounding use. A log analyzer must not offer it as a log-match pattern.
Full UA: none exists — Google-Extended is a robots.txt-only token with no user-agent string; any log line carrying it is a spoofer.
How to verify a hit is really Google-Extended
Verification method (July 2026): n/a — robots.txt control token only
User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to Google — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.
Allow or block in robots.txt
Add the token to robots.txt — this is the only place Google-Extended does anything:
# Block Google-Extended site-wide
User-agent: Google-Extended
Disallow: /
# Permit AI use of crawled content
User-agent: Google-Extended
Allow: /
For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.
Vendor documentation: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
Frequently asked questions
What is Google-Extended?
Google-Extended is a robots.txt control token from Google — it has no user agent and never appears in access logs; it controls AI-training use of content crawled by Google's regular crawlers.
Does Google-Extended respect robots.txt?
Yes — Google documents robots.txt compliance for Google-Extended (vendor docs reviewed July 2026).
How do I block Google-Extended?
Add 'User-agent: Google-Extended' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.