The AI Crawler robots.txt Guide: GPTBot, ClaudeBot, PerplexityBot, and the Rest

By · Published

The AI Crawler robots.txt Guide: GPTBot, ClaudeBot, PerplexityBot, and the Rest

Most robots.txt files I open were written for a web with a handful of crawlers that mattered. Now every AI company runs two or three bots, each with a different job, and the names are close enough that people block the wrong one. A common mistake: someone blocks GPTBot to stay out of training, then wonders why they never show up in ChatGPT search. GPTBot was never the search crawler.

Below are the tokens that matter right now, what each one does, and recipes you can paste in. Everything is checked against the providers’ own docs, linked inline.

Three kinds of AI bots

AI traffic falls into three buckets, and blocking each one has very different consequences.

  1. Training crawlers. They collect pages that may end up in a future model’s training data. Blocking them keeps your future content out of that pipeline. It does not remove anything already collected.
  2. Search and indexing crawlers. They build the index an AI assistant searches when it answers questions and cites sources. Blocking them is the AI equivalent of blocking Googlebot. You stop getting cited.
  3. User-triggered fetchers. A person pastes your URL into a chat or asks a question that needs a live page, and the assistant fetches it on their behalf. Several providers say plainly that robots.txt may not apply to these, because a human asked.

Most people want to block bucket one and allow buckets two and three. Some want to block everything. Both are reasonable. Just don’t do it by accident.

The user-agent tokens

OpenAI

From OpenAI’s crawler docs:

  • GPTBot - training. “Disallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models.”
  • OAI-SearchBot - search. It surfaces sites in ChatGPT’s search features. If you block it, you can still show up as a navigational link, but not in search answers.
  • ChatGPT-User - user-triggered fetches. OpenAI says it isn’t used for automatic crawling, isn’t used to decide search inclusion, and “robots.txt rules may not apply.”

Two details worth knowing. If you allow both GPTBot and OAI-SearchBot, OpenAI may reuse one crawl for both purposes. And OpenAI says it can take about 24 hours for its systems to pick up a robots.txt change. OpenAI also lists OAI-AdsBot, which does ad safety checks and ignores robots.txt, but that one only matters if you run ads on their platform.

Anthropic

From Anthropic’s crawler article:

  • ClaudeBot - training. Blocking it signals that your future content should be excluded from training datasets.
  • Claude-SearchBot - search indexing, used to improve Claude’s search results.
  • Claude-User - fetches pages when a Claude user asks a question that needs them.

Anthropic says its bots honor robots.txt, and it documents blocking each of the three separately. It also supports Crawl-delay, which a lot of crawlers ignore.

Perplexity

From Perplexity’s bot docs:

  • PerplexityBot - search. Perplexity says it surfaces and links sites in results and “is not used to crawl content for AI foundation models.”
  • Perplexity-User - user-triggered. Perplexity states this fetcher “generally ignores robots.txt rules.”

Perplexity doesn’t list a training crawler, so there isn’t a training token to block.

Google

From Google’s common crawlers docs:

  • Google-Extended - controls whether content Google crawls can be used to train future Gemini models and for grounding in the Gemini Apps and the Vertex AI API. It has no HTTP user agent of its own. Googlebot does the crawling, and the token acts as a permission flag. Google says it doesn’t affect Search inclusion or ranking.

This is where people get burned. Blocking Google-Extended does not remove you from AI Overviews or AI Mode, because those are Search features. Google added a separate Search generative AI control in Search Console settings for that. It’s a Search Console toggle, not a robots.txt rule, and Google says it doesn’t affect regular Search.

Google also lists user-triggered fetchers like Google-Agent and Google-GeminiNotebook, which “generally ignore robots.txt rules.”

Apple

From Apple’s Applebot docs:

  • Applebot - powers search in Spotlight, Siri, and Safari.
  • Applebot-Extended - controls whether Applebot’s data can be used to train Apple’s generative models. Like Google-Extended, it “does not crawl webpages.” Apple says pages that disallow Applebot-Extended can still show up in search results.

Common Crawl

CCBot builds the Common Crawl public dataset. Common Crawl isn’t an AI company, but its archive is widely used in LLM training data. If your goal is “stay out of training,” this one belongs on the list.

Meta

From Meta’s crawler docs:

  • meta-externalagent - crawls for use cases like training foundation models and improving products by indexing content.
  • meta-webindexer - improves Meta AI search results.
  • meta-externalfetcher - user-requested fetches, and Meta says it may bypass robots.txt.
  • facebookexternalhit - link previews when someone shares your URL. Block this and your shared links look broken.

The grouping rule that breaks everything

Before you paste anything, understand how groups work. Under RFC 9309 and Google’s implementation, a crawler picks the single most specific group that matches its name and ignores the rest. A specific group and the * group are not combined.

So if your file looks like this:

User-agent: *
Disallow: /admin/
Disallow: /cart/

User-agent: GPTBot
Allow: /

GPTBot follows only its own group, and /admin/ and /cart/ are wide open to it. Any path you want kept away from every bot has to be repeated in each named group. This is an easy bug to ship and a hard one to notice.

Token matching is case-insensitive, and the order of groups in the file doesn’t matter.

Recipe 1: allow everything

If you want maximum AI visibility, you don’t need to name anyone. A wildcard group covers them:

User-agent: *
Disallow: /admin/
Disallow: /cart/

Sitemap: https://www.example.com/sitemap.xml

Recipe 2: block training, allow search and user fetches

This is the one most businesses want. You stay citable in ChatGPT, Claude, Perplexity, and Google, but opt out of future training.

# Training and dataset crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
Disallow: /

# Everyone else, including OAI-SearchBot, Claude-SearchBot,
# PerplexityBot, Claude-User, ChatGPT-User, Googlebot, Applebot
User-agent: *
Disallow: /admin/
Disallow: /cart/

Sitemap: https://www.example.com/sitemap.xml

Stacking several User-agent lines over one set of rules is valid under RFC 9309. If you’d rather not trust every parser on that, write a separate group per bot. It’s uglier but unambiguous.

A judgment call: meta-externalagent covers both training and product indexing, so blocking it may cost you some visibility in Meta’s products. Decide based on whether that traffic matters to you.

Recipe 3: block all AI bots

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: meta-externalfetcher
Disallow: /

User-agent: *
Disallow: /admin/
Disallow: /cart/

Understand what you’re giving up. You won’t be cited in AI answers from those providers, and you won’t get the referral traffic that comes with citations. Also, several of the user-triggered fetchers in that list say they may ignore robots.txt, so this file alone won’t stop them. If you truly need them gone, that’s a firewall job.

Notice what’s missing: Googlebot and Applebot. Blocking those kills your regular search presence. Google-Extended and Applebot-Extended are the AI-specific levers.

robots.txt is a request, not a lock

RFC 9309 says it directly: robots.txt rules are not a form of access authorization. Nothing enforces them. Reputable crawlers follow them because it’s a norm and because the providers have publicly committed to it.

In 2025, Cloudflare accused Perplexity of using undeclared crawlers with generic browser user agents to get around blocks, and delisted it as a verified bot. Perplexity disputed the analysis. Either way, a user-agent string is a claim anyone can make, and robots.txt only works on bots that choose to honor it.

If content must not be fetched, put it behind authentication or block it at the edge.

Verify bots by IP, not by user agent

Anyone can send User-Agent: GPTBot with curl. Before you trust or block traffic based on that string, check the source IP. Most of the big providers now publish their ranges:

For Google, the documented check is a reverse lookup followed by a forward lookup that must match:

host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1

For the providers that publish JSON, a short script against your access logs does the job:

import ipaddress, json, urllib.request

def load_ranges(url):
    data = json.load(urllib.request.urlopen(url))
    nets = []
    for p in data.get("prefixes", []):
        cidr = p.get("ipv4Prefix") or p.get("ipv6Prefix")
        if cidr:
            nets.append(ipaddress.ip_network(cidr))
    return nets

gptbot = load_ranges("https://openai.com/gptbot.json")

def is_real_gptbot(ip):
    addr = ipaddress.ip_address(ip)
    return any(addr in net for net in gptbot)

print(is_real_gptbot("203.0.113.9"))  # a documentation IP, should be False

These lists change, so fetch them on a schedule instead of hardcoding them.

How to test your robots.txt

Read what you’re actually serving. curl -s https://www.example.com/robots.txt. Staging files get deployed to production, and CMS plugins rewrite the file quietly.

Check that your edge isn’t overriding you. CDNs and WAFs, Cloudflare included, have AI bot blocking settings. A robots.txt that allows OAI-SearchBot does nothing if your firewall returns a 403 first:

curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot" https://www.example.com/

A spoofed user agent won’t match a verified-bot rule, so this won’t catch everything, but it will catch a blanket block on the string.

Use Google’s tooling for Google. The robots.txt report in Search Console shows what Google fetched and any parse errors. Google also open-sourced its robots.txt parser if you want the same matching logic locally.

Be careful with Python’s urllib.robotparser. It’s handy for quick checks, but it applies the first matching rule in file order instead of the longest match like Google and RFC 9309 do. If your file mixes Allow and Disallow on overlapping paths, its answers can differ from what real crawlers do.

Test real URLs, not just the homepage. Robots rules interact with redirects, canonicals, and sitemaps. Pull a list of your important URLs from the sitemap and check each one against each AI group you wrote. The grouping mistake above usually shows up on the second or third path you test, not the first.

AI crawlers bring two more problems with them. Most of them don’t run JavaScript, which I cover in Can AI crawlers see your JavaScript content?. And people keep asking whether an llms.txt file belongs next to robots.txt, which is a separate question with a less exciting answer.

The fundamentals haven’t changed much. Google says SEO for AI is the same as SEO for search, and the bots can only cite pages they’re allowed to fetch.

Letting AI bots in only helps if the pages they fetch are in good shape. Black SEO Analyzer runs locally on Windows, macOS, and Linux, with 16 analysis modules and 300+ issue types. Try it free for 14 days, with every feature, no page limit, and no credit card.

-Sethers

Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

Download Free Trial

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.