Robots.txt Tester

Check whether a URL is allowed or blocked by robots.txt for Googlebot, Bingbot, GPTBot or any other crawler. We fetch the site's robots.txt (or use one you paste), apply Google's matching rules, and show the exact line that decided it.
Paste robots.txt instead of fetching it (optional)

Useful for testing changes before you publish them. With pasted content, the URL can be a path like /private/page. Limit 200 KB.

How to check if a page is blocked by robots.txt

Enter the full URL of the page and pick a crawler. The tester fetches /robots.txt from that URL's host, finds the group of rules for that crawler, and applies the longest matching rule, the same way Google does. You get Allowed or Blocked, the line that decided it, and the verdict for other common crawlers.

Google used to have a robots.txt Tester inside Search Console. It was retired in December 2023 and replaced by the robots.txt report, which shows the robots.txt files Google fetched and any parse errors, but doesn't let you test a URL against the rules. The URL Inspection tool in Search Console still tells you if Google was blocked from crawling a specific URL on your verified site.

To test changes before you publish them, paste the new file into the "Paste robots.txt" box. Nothing is fetched and you can enter just a path, such as /checkout/step-1.

What robots.txt does, and what it doesn't

robots.txt controls crawling. It tells well-behaved crawlers which URLs they may request. It does not control indexing. If other pages link to a blocked URL, Google can still index it and show it in results, usually with no description, because it never read the page.

To keep a page out of search results, use a noindex robots meta tag or an X-Robots-Tag: noindex HTTP header, and leave the page crawlable. If you also block it in robots.txt, Google can't fetch the page, never sees the noindex, and the URL can stay indexed. Google has not supported noindex lines inside robots.txt since 2019.

robots.txt is also not access control. Anyone can read the file, and bad bots ignore it. Protect private content with authentication.

Where the file must live

A crawler only looks for robots.txt at the root of a host: https://www.example.com/robots.txt. Rules apply only to the exact protocol, host and port the file is served from. https://example.com/robots.txt doesn't cover https://www.example.com/ or https://shop.example.com/, and a file at /blog/robots.txt is ignored. The file must be UTF-8 plain text, and Google reads only the first 500 KiB.

How the matching rules work

A file is made of groups. Each group starts with one or more User-agent lines followed by Allow and Disallow rules:

User-agent: Googlebot
User-agent: Bingbot
Disallow: /search
Allow: /search/about

User-agent: *
Disallow: /admin/

Sitemap: https://www.example.com/sitemap.xml

1. Pick the group

A crawler uses the group whose user agent matches its product token most specifically. Matching is case-insensitive, and anything after the token is ignored, so googlebot/1.2 means googlebot. If several groups name the same crawler, their rules are combined. Only if no group names the crawler does it use the * group. Groups don't stack: in the example above, Googlebot follows only the first group, so /admin/ is allowed for Googlebot.

Some crawlers have fallbacks. Googlebot-Image follows a Googlebot group if there's no Googlebot-Image group, and Apple says Applebot follows Googlebot rules when Applebot isn't mentioned. The table in the results reflects both.

2. Pick the rule

Rules are matched against the URL's path and query string, starting at the first /. Paths are case-sensitive. Among the rules that match, the one with the longest path wins. If an Allow and a Disallow rule are the same length, Allow wins. The order of lines doesn't matter. These examples come from Google's robots.txt documentation:

URL pathRulesResult
/pageAllow: /p
Disallow: /
Allowed: /p is longer
/folder/pageAllow: /folder
Disallow: /folder
Allowed: same length, Allow wins
/page.htmAllow: /page
Disallow: /*.htm
Blocked: /*.htm is longer
/Allow: /$
Disallow: /
Allowed: /$ is longer

An empty Disallow: matches nothing, so a group with only Disallow: allows everything.

3. Wildcards

* matches any sequence of characters, including none. $ at the end of a rule means the URL must end there. Without $, every rule is a prefix match.

Rule pathMatchesDoesn't match
/fish/fish, /fish.html, /fishheads/Fish.asp, /catfish
/fish//fish/, /fish/?id=anything/fish, /fish.html
/*.php/index.php, /folder/filename.php/
/*.php$/filename.php/filename.php?parameters
/*?any URL with a query string/page

Non-ASCII characters are compared in percent-encoded form, so Disallow: /foo/bar/ツ and Disallow: /foo/bar/%E3%83%84 are the same rule.

Sitemap and other lines

Sitemap: lines don't belong to any group and apply to every crawler. The value must be a full URL including https:// and the host. Use the XML Sitemap Checker to validate the files they point to. Google ignores Crawl-delay, although Bing honors it. Lines Google doesn't recognize are ignored, and this tester lists them so you can spot typos.

What happens when robots.txt returns an error

ResponseHow Google treats it
2xxParses the file as served.
3xxFollows at least five redirect hops, then treats it as a 404.
4xx except 429As if there were no robots.txt: everything may be crawled.
5xx, 429, timeouts, DNS or connection errorsThe whole site is treated as disallowed. For the first 12 hours Google stops crawling and keeps retrying, then uses its last cached copy for up to 30 days. After that, if the site is otherwise available, it crawls as if there were no robots.txt.

A robots.txt that times out or throws a 500 can stop Google from crawling your whole site, even if every page works. Google normally caches robots.txt for up to 24 hours, so changes aren't picked up instantly.

Testing AI crawlers

AI companies use their own user agent tokens, such as GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot and CCBot (Common Crawl). If your file has no group for them, they follow the * group. A Disallow: / under User-agent: GPTBot doesn't affect Googlebot, and blocking Googlebot doesn't block GPTBot. Some tokens, like Google-Extended and Applebot-Extended, control how content is used for AI features and training, and don't crawl pages on their own. Test them with the custom field. Our AI crawler robots.txt guide lists the tokens and what each one controls.

Common robots.txt mistakes

Checking robots.txt during a full crawl

This tester checks one URL at a time. bseoa, our desktop crawler, reads each host's robots.txt when it crawls your site, skips the URLs it disallows, and records each one with a "Not crawled: blocked by robots.txt" warning, so you can see which discovered URLs your rules block across the whole site. When you audit your own site and want those pages analyzed anyway, add --ignore-robots to the CLI crawl.

Check every page, not just one

These tools look at one URL at a time. bseoa crawls your whole site on your own machine and checks every page for 300+ technical SEO issues across 16 analysis modules.

Start the 14-day free trial

More free SEO tools