The AI Crawler robots.txt Guide: GPTBot, ClaudeBot, PerplexityBot, and the Rest
By Seth Black · Published
Most robots.txt files I open were written for a web with a handful of crawlers that mattered. Now every AI company runs two or three bots, each with a different job, and the names are close enough that people block the wrong one. A common mistake: someone blocks GPTBot to stay out of training, then wonders why they never show up in ChatGPT search. GPTBot was never the search crawler.
Below are the tokens that matter right now, what each one does, and recipes you can paste in. Everything is checked against the providers’ own docs, linked inline.
Three kinds of AI bots
AI traffic falls into three buckets, and blocking each one has very different consequences.
- Training crawlers. They collect pages that may end up in a future model’s training data. Blocking them keeps your future content out of that pipeline. It does not remove anything already collected.
- Search and indexing crawlers. They build the index an AI assistant searches when it answers questions and cites sources. Blocking them is the AI equivalent of blocking Googlebot. You stop getting cited.
- User-triggered fetchers. A person pastes your URL into a chat or asks a question that needs a live page, and the assistant fetches it on their behalf. Several providers say plainly that robots.txt may not apply to these, because a human asked.
Most people want to block bucket one and allow buckets two and three. Some want to block everything. Both are reasonable. Just don’t do it by accident.
The user-agent tokens
OpenAI
From OpenAI’s crawler docs:
- GPTBot - training. “Disallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models.”
- OAI-SearchBot - search. It surfaces sites in ChatGPT’s search features. If you block it, you can still show up as a navigational link, but not in search answers.
- ChatGPT-User - user-triggered fetches. OpenAI says it isn’t used for automatic crawling, isn’t used to decide search inclusion, and “robots.txt rules may not apply.”
Two details worth knowing. If you allow both GPTBot and OAI-SearchBot, OpenAI may reuse one crawl for both purposes. And OpenAI says it can take about 24 hours for its systems to pick up a robots.txt change. OpenAI also lists OAI-AdsBot, which does ad safety checks and ignores robots.txt, but that one only matters if you run ads on their platform.
Anthropic
From Anthropic’s crawler article:
- ClaudeBot - training. Blocking it signals that your future content should be excluded from training datasets.
- Claude-SearchBot - search indexing, used to improve Claude’s search results.
- Claude-User - fetches pages when a Claude user asks a question that needs them.
Anthropic says its bots honor robots.txt, and it documents blocking each of the three separately. It also supports Crawl-delay, which a lot of crawlers ignore.
Perplexity
From Perplexity’s bot docs:
- PerplexityBot - search. Perplexity says it surfaces and links sites in results and “is not used to crawl content for AI foundation models.”
- Perplexity-User - user-triggered. Perplexity states this fetcher “generally ignores robots.txt rules.”
Perplexity doesn’t list a training crawler, so there isn’t a training token to block.
From Google’s common crawlers docs:
- Google-Extended - controls whether content Google crawls can be used to train future Gemini models and for grounding in the Gemini Apps and the Vertex AI API. It has no HTTP user agent of its own. Googlebot does the crawling, and the token acts as a permission flag. Google says it doesn’t affect Search inclusion or ranking.
This is where people get burned. Blocking Google-Extended does not remove you from AI Overviews or AI Mode, because those are Search features. Google added a separate Search generative AI control in Search Console settings for that. It’s a Search Console toggle, not a robots.txt rule, and Google says it doesn’t affect regular Search.
Google also lists user-triggered fetchers like Google-Agent and Google-GeminiNotebook, which “generally ignore robots.txt rules.”
Apple
From Apple’s Applebot docs:
- Applebot - powers search in Spotlight, Siri, and Safari.
- Applebot-Extended - controls whether Applebot’s data can be used to train Apple’s generative models. Like Google-Extended, it “does not crawl webpages.” Apple says pages that disallow Applebot-Extended can still show up in search results.
Common Crawl
CCBot builds the Common Crawl public dataset. Common Crawl isn’t an AI company, but its archive is widely used in LLM training data. If your goal is “stay out of training,” this one belongs on the list.
Meta
From Meta’s crawler docs:
- meta-externalagent - crawls for use cases like training foundation models and improving products by indexing content.
- meta-webindexer - improves Meta AI search results.
- meta-externalfetcher - user-requested fetches, and Meta says it may bypass robots.txt.
- facebookexternalhit - link previews when someone shares your URL. Block this and your shared links look broken.
The grouping rule that breaks everything
Before you paste anything, understand how groups work. Under RFC 9309 and Google’s implementation, a crawler picks the single most specific group that matches its name and ignores the rest. A specific group and the * group are not combined.
So if your file looks like this:
User-agent: *
Disallow: /admin/
Disallow: /cart/
User-agent: GPTBot
Allow: /
GPTBot follows only its own group, and /admin/ and /cart/ are wide open to it. Any path you want kept away from every bot has to be repeated in each named group. This is an easy bug to ship and a hard one to notice.
Token matching is case-insensitive, and the order of groups in the file doesn’t matter.
Recipe 1: allow everything
If you want maximum AI visibility, you don’t need to name anyone. A wildcard group covers them:
User-agent: *
Disallow: /admin/
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xml
Recipe 2: block training, allow search and user fetches
This is the one most businesses want. You stay citable in ChatGPT, Claude, Perplexity, and Google, but opt out of future training.
# Training and dataset crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
Disallow: /
# Everyone else, including OAI-SearchBot, Claude-SearchBot,
# PerplexityBot, Claude-User, ChatGPT-User, Googlebot, Applebot
User-agent: *
Disallow: /admin/
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xml
Stacking several User-agent lines over one set of rules is valid under RFC 9309. If you’d rather not trust every parser on that, write a separate group per bot. It’s uglier but unambiguous.
A judgment call: meta-externalagent covers both training and product indexing, so blocking it may cost you some visibility in Meta’s products. Decide based on whether that traffic matters to you.
Recipe 3: block all AI bots
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: meta-externalfetcher
Disallow: /
User-agent: *
Disallow: /admin/
Disallow: /cart/
Understand what you’re giving up. You won’t be cited in AI answers from those providers, and you won’t get the referral traffic that comes with citations. Also, several of the user-triggered fetchers in that list say they may ignore robots.txt, so this file alone won’t stop them. If you truly need them gone, that’s a firewall job.
Notice what’s missing: Googlebot and Applebot. Blocking those kills your regular search presence. Google-Extended and Applebot-Extended are the AI-specific levers.
robots.txt is a request, not a lock
RFC 9309 says it directly: robots.txt rules are not a form of access authorization. Nothing enforces them. Reputable crawlers follow them because it’s a norm and because the providers have publicly committed to it.
In 2025, Cloudflare accused Perplexity of using undeclared crawlers with generic browser user agents to get around blocks, and delisted it as a verified bot. Perplexity disputed the analysis. Either way, a user-agent string is a claim anyone can make, and robots.txt only works on bots that choose to honor it.
If content must not be fetched, put it behind authentication or block it at the edge.
Verify bots by IP, not by user agent
Anyone can send User-Agent: GPTBot with curl. Before you trust or block traffic based on that string, check the source IP. Most of the big providers now publish their ranges:
- OpenAI: gptbot.json, searchbot.json, chatgpt-user.json
- Anthropic: bots.json
- Perplexity: perplexitybot.json, perplexity-user.json
- Common Crawl: ccbot.json, plus reverse DNS under
crawl.commoncrawl.org - Apple: applebot.json, plus reverse DNS under
applebot.apple.com - Google: several JSON files, plus reverse DNS
For Google, the documented check is a reverse lookup followed by a forward lookup that must match:
host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1
For the providers that publish JSON, a short script against your access logs does the job:
import ipaddress, json, urllib.request
def load_ranges(url):
data = json.load(urllib.request.urlopen(url))
nets = []
for p in data.get("prefixes", []):
cidr = p.get("ipv4Prefix") or p.get("ipv6Prefix")
if cidr:
nets.append(ipaddress.ip_network(cidr))
return nets
gptbot = load_ranges("https://openai.com/gptbot.json")
def is_real_gptbot(ip):
addr = ipaddress.ip_address(ip)
return any(addr in net for net in gptbot)
print(is_real_gptbot("203.0.113.9")) # a documentation IP, should be False
These lists change, so fetch them on a schedule instead of hardcoding them.
How to test your robots.txt
Read what you’re actually serving. curl -s https://www.example.com/robots.txt. Staging files get deployed to production, and CMS plugins rewrite the file quietly.
Check that your edge isn’t overriding you. CDNs and WAFs, Cloudflare included, have AI bot blocking settings. A robots.txt that allows OAI-SearchBot does nothing if your firewall returns a 403 first:
curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot" https://www.example.com/
A spoofed user agent won’t match a verified-bot rule, so this won’t catch everything, but it will catch a blanket block on the string.
Use Google’s tooling for Google. The robots.txt report in Search Console shows what Google fetched and any parse errors. Google also open-sourced its robots.txt parser if you want the same matching logic locally.
Be careful with Python’s urllib.robotparser. It’s handy for quick checks, but it applies the first matching rule in file order instead of the longest match like Google and RFC 9309 do. If your file mixes Allow and Disallow on overlapping paths, its answers can differ from what real crawlers do.
Test real URLs, not just the homepage. Robots rules interact with redirects, canonicals, and sitemaps. Pull a list of your important URLs from the sitemap and check each one against each AI group you wrote. The grouping mistake above usually shows up on the second or third path you test, not the first.
Related reading
AI crawlers bring two more problems with them. Most of them don’t run JavaScript, which I cover in Can AI crawlers see your JavaScript content?. And people keep asking whether an llms.txt file belongs next to robots.txt, which is a separate question with a less exciting answer.
The fundamentals haven’t changed much. Google says SEO for AI is the same as SEO for search, and the bots can only cite pages they’re allowed to fetch.
Letting AI bots in only helps if the pages they fetch are in good shape. Black SEO Analyzer runs locally on Windows, macOS, and Linux, with 16 analysis modules and 300+ issue types. Try it free for 14 days, with every feature, no page limit, and no credit card.
-Sethers