What it does
GPTBot is OpenAI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Full UA: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot. robots.txt fetches may carry an extra 'robots.txt' marker in the UA. Robots compliance field-confirmed on our own properties (a production site we operate 2026-07-01: hammering stopped within one crawl cycle of a Disallow).
How to verify a hit is really GPTBot
Verification method (July 2026): Published IP list: https://openai.com/gptbot.json (fetched live 2026-07-19)
User-agent strings are freely spoofed, so a UA match alone proves nothing. Confirm the source IP belongs to OpenAI — published IP-range JSON where available, otherwise RDAP/rDNS on the IP. Full method: verify by IP.
Allow or block in robots.txt
Match the exact token GPTBot in robots.txt:
# Block GPTBot site-wide
User-agent: GPTBot
Disallow: /
# Explicitly allow GPTBot
User-agent: GPTBot
Allow: /
For the allow-search-block-training combined pattern (and the Allow-directive gotcha that silently kills carve-outs), see the robots.txt guide.
Vendor documentation: https://platform.openai.com/docs/bots
Frequently asked questions
What is GPTBot?
GPTBot is OpenAI's training crawler. Content is copied into datasets used to train foundation models. Blocking costs you nothing in traffic today; allowing is a donation of your content to model weights with no citation or referral in return.
Does GPTBot respect robots.txt?
Yes — OpenAI documents robots.txt compliance for GPTBot (vendor docs reviewed July 2026).
How do I block GPTBot?
Add 'User-agent: GPTBot' followed by 'Disallow: /' to your robots.txt. Use the exact token — substring variants of other tokens will not match.