---
title: "Robots.txt Tester: Check If a URL Is Blocked"
description: "Free robots.txt tester. Fetch or paste a robots.txt file, test any URL against Googlebot or another user agent, and see exactly which rule allows or blocks it."
image: "https://www.blackseoanalyzer.com/static/images/black-seo-analyzer-og-image.png"
canonical: "https://www.blackseoanalyzer.com/en/tools/robots-txt-tester"
language: "en"
---

# Robots.txt Tester

Check whether a URL is allowed or blocked by robots.txt for Googlebot, Bingbot, GPTBot or any other crawler. We fetch the site's robots.txt (or use one you paste), apply Google's matching rules, and show the exact line that decided it.

## How to check if a page is blocked by robots.txt

Enter the full URL of the page and pick a crawler. The tester fetches `/robots.txt` from that URL's host, finds the group of rules for that crawler, and applies the longest matching rule, the same way Google does. You get Allowed or Blocked, the line that decided it, and the verdict for other common crawlers.

Google used to have a robots.txt Tester inside Search Console. It was retired in December 2023 and replaced by the [robots.txt report](https://support.google.com/webmasters/answer/6062598), which shows the robots.txt files Google fetched and any parse errors, but doesn't let you test a URL against the rules. The URL Inspection tool in Search Console still tells you if Google was blocked from crawling a specific URL on your verified site.

To test changes before you publish them, paste the new file into the "Paste robots.txt" box. Nothing is fetched and you can enter just a path, such as `/checkout/step-1`.

## What robots.txt does, and what it doesn't

robots.txt controls **crawling**. It tells well-behaved crawlers which URLs they may request. It does not control **indexing**. If other pages link to a blocked URL, Google can still index it and show it in results, usually with no description, because it never read the page.

To keep a page out of search results, use a `noindex` robots meta tag or an `X-Robots-Tag: noindex` HTTP header, and leave the page crawlable. If you also block it in robots.txt, Google can't fetch the page, never sees the `noindex`, and the URL can stay indexed. Google has not supported `noindex` lines inside robots.txt since 2019.

robots.txt is also not access control. Anyone can read the file, and bad bots ignore it. Protect private content with authentication.

## Where the file must live

A crawler only looks for robots.txt at the root of a host: `https://www.example.com/robots.txt`. Rules apply only to the exact protocol, host and port the file is served from. `https://example.com/robots.txt` doesn't cover `https://www.example.com/` or `https://shop.example.com/`, and a file at `/blog/robots.txt` is ignored. The file must be UTF-8 plain text, and Google reads only the first 500 KiB.

## How the matching rules work

A file is made of groups. Each group starts with one or more `User-agent` lines followed by `Allow` and `Disallow` rules:

```
User-agent: Googlebot
User-agent: Bingbot
Disallow: /search
Allow: /search/about

User-agent: *
Disallow: /admin/

Sitemap: https://www.example.com/sitemap.xml
```

### 1. Pick the group

A crawler uses the group whose user agent matches its product token most specifically. Matching is case-insensitive, and anything after the token is ignored, so `googlebot/1.2` means `googlebot`. If several groups name the same crawler, their rules are combined. Only if no group names the crawler does it use the `*` group. Groups don't stack: in the example above, Googlebot follows only the first group, so `/admin/` is allowed for Googlebot.

Some crawlers have fallbacks. Googlebot-Image follows a `Googlebot` group if there's no `Googlebot-Image` group, and Apple says Applebot follows Googlebot rules when Applebot isn't mentioned. The table in the results reflects both.

### 2. Pick the rule

Rules are matched against the URL's path and query string, starting at the first `/`. Paths are case-sensitive. Among the rules that match, the one with the **longest path** wins. If an Allow and a Disallow rule are the same length, **Allow wins**. The order of lines doesn't matter. These examples come from [Google's robots.txt documentation](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt):

| URL path | Rules | Result |
| --- | --- | --- |
| `/page` | `Allow: /p` `Disallow: /` | Allowed: `/p` is longer |
| `/folder/page` | `Allow: /folder` `Disallow: /folder` | Allowed: same length, Allow wins |
| `/page.htm` | `Allow: /page` `Disallow: /*.htm` | Blocked: `/*.htm` is longer |
| `/` | `Allow: /$` `Disallow: /` | Allowed: `/$` is longer |

An empty `Disallow:` matches nothing, so a group with only `Disallow:` allows everything.

### 3. Wildcards

`*` matches any sequence of characters, including none. `$` at the end of a rule means the URL must end there. Without `$`, every rule is a prefix match.

| Rule path | Matches | Doesn't match |
| --- | --- | --- |
| `/fish` | `/fish`, `/fish.html`, `/fishheads` | `/Fish.asp`, `/catfish` |
| `/fish/` | `/fish/`, `/fish/?id=anything` | `/fish`, `/fish.html` |
| `/*.php` | `/index.php`, `/folder/filename.php` | `/` |
| `/*.php$` | `/filename.php` | `/filename.php?parameters` |
| `/*?` | any URL with a query string | `/page` |

Non-ASCII characters are compared in percent-encoded form, so `Disallow: /foo/bar/ツ` and `Disallow: /foo/bar/%E3%83%84` are the same rule.

### Sitemap and other lines

`Sitemap:` lines don't belong to any group and apply to every crawler. The value must be a full URL including `https://` and the host. Use the [XML Sitemap Checker](https://www.blackseoanalyzer.com/en/tools/xml-sitemap-checker) to validate the files they point to. Google ignores `Crawl-delay`, although Bing honors it. Lines Google doesn't recognize are ignored, and this tester lists them so you can spot typos.

## What happens when robots.txt returns an error

| Response | How Google treats it |
| --- | --- |
| 2xx | Parses the file as served. |
| 3xx | Follows at least five redirect hops, then treats it as a 404. |
| 4xx except 429 | As if there were no robots.txt: everything may be crawled. |
| 5xx, 429, timeouts, DNS or connection errors | The whole site is treated as disallowed. For the first 12 hours Google stops crawling and keeps retrying, then uses its last cached copy for up to 30 days. After that, if the site is otherwise available, it crawls as if there were no robots.txt. |

A robots.txt that times out or throws a 500 can stop Google from crawling your whole site, even if every page works. Google normally caches robots.txt for up to 24 hours, so changes aren't picked up instantly.

## Testing AI crawlers

AI companies use their own user agent tokens, such as `GPTBot` (OpenAI), `ClaudeBot` (Anthropic), `PerplexityBot` and `CCBot` (Common Crawl). If your file has no group for them, they follow the `*` group. A `Disallow: /` under `User-agent: GPTBot` doesn't affect Googlebot, and blocking Googlebot doesn't block GPTBot. Some tokens, like `Google-Extended` and `Applebot-Extended`, control how content is used for AI features and training, and don't crawl pages on their own. Test them with the custom field. Our [AI crawler robots.txt guide](https://www.blackseoanalyzer.com/en/blog/ai-crawler-robots-txt-guide) lists the tokens and what each one controls.

## Common robots.txt mistakes

- **`Disallow: /` left over from staging.** A site launched with the staging robots.txt blocks every crawler from everything. Check the file after every deploy.
- **Blocking CSS and JavaScript.** Google renders pages. If it can't fetch your stylesheets and scripts, it may not see your content or layout correctly. Don't disallow `/assets/`, `/static/` or similar folders that pages need.
- **Wrong case.** `Disallow: /Admin/` doesn't block `/admin/`. Paths are case-sensitive; user agent names are not.
- **Expecting groups to combine with `*`.** Adding a `User-agent: Googlebot` group for one rule means Googlebot no longer follows anything in the `*` group. Repeat the rules you still want.
- **Using robots.txt to remove pages from Google.** Blocking a URL stops crawling, not indexing. Use `noindex` and let the page be crawled.
- **Rules without a leading slash.** `Disallow: private` never matches, because every URL path starts with `/`.
- **Relative sitemap URLs.** `Sitemap: /sitemap.xml` is invalid. Use the full URL.

## Checking robots.txt during a full crawl

This tester checks one URL at a time. bseoa, our desktop crawler, reads each host's robots.txt when it crawls your site, skips the URLs it disallows, and records each one with a "Not crawled: blocked by robots.txt" warning, so you can see which discovered URLs your rules block across the whole site. When you audit your own site and want those pages analyzed anyway, add `--ignore-robots` to the CLI crawl.

## Check every page, not just one

These tools look at one URL at a time. bseoa crawls your whole site on your own machine and checks every page for 300+ technical SEO issues across 16 analysis modules.

[Start the 14-day free trial](https://www.blackseoanalyzer.com/en/free-trial)

## More free SEO tools

- [Title & Meta Description Length Checker](https://www.blackseoanalyzer.com/en/tools/meta-description-length-checker)
Type or paste a title and meta description and see, as you type, whether Google or Bing will cut them off, measured in pixels rather than characters.

- [XML Sitemap Checker](https://www.blackseoanalyzer.com/en/tools/xml-sitemap-checker)
Find a site's sitemap from its domain, then validate the XML, the URLs in it, lastmod dates and sitemap index files.

- [Schema Markup Validator](https://www.blackseoanalyzer.com/en/tools/schema-markup-validator)
Paste JSON-LD or enter a URL. Get JSON syntax errors with line numbers and a check of the properties Google requires for rich results.

- [Hreflang Checker](https://www.blackseoanalyzer.com/en/tools/hreflang-checker)
Validate a page's hreflang annotations, then fetch each alternate to confirm it links back, returns 200 and isn't canonicalized elsewhere.

- [Canonical Tag Checker](https://www.blackseoanalyzer.com/en/tools/canonical-tag-checker)
See which URL a page declares as canonical, in HTML or HTTP headers, and whether that target actually resolves to an indexable page.

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://www.blackseoanalyzer.com/#organization","name":"Fiscus Technology, LLC","url":"https://www.blackseoanalyzer.com/","logo":"https://www.blackseoanalyzer.com/static/images/black-seo-analyzer.png"},{"@type":"WebSite","@id":"https://www.blackseoanalyzer.com/#website","name":"bseoa","url":"https://www.blackseoanalyzer.com/","publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"}},{"@type":"WebApplication","@id":"https://www.blackseoanalyzer.com/en/tools/robots-txt-tester#app","name":"Robots.txt Tester","url":"https://www.blackseoanalyzer.com/en/tools/robots-txt-tester","description":"Test whether a URL is allowed or blocked for Googlebot, Bingbot, GPTBot or any user agent, and see the exact rule that decided it.","applicationCategory":"DeveloperApplication","operatingSystem":"Any (web browser)","isAccessibleForFree":true,"offers":{"@type":"Offer","price":"0","priceCurrency":"USD"},"publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"}},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://www.blackseoanalyzer.com/"},{"@type":"ListItem","position":2,"name":"Free SEO Tools","item":"https://www.blackseoanalyzer.com/en/tools"},{"@type":"ListItem","position":3,"name":"Robots.txt Tester","item":"https://www.blackseoanalyzer.com/en/tools/robots-txt-tester"}]}]}
```
