---
title: "How to See All Pages on a Website (7 Ways That Work)"
description: "See every page on a website using the sitemap, a site: search, Search Console, a crawler, your CMS, analytics or the Wayback Machine, and find orphan pages."
image: "https://www.blackseoanalyzer.com/static/images/black-seo-analyzer-og-image.png"
canonical: "https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website"
language: "en"
---

# How to See All Pages on a Website

By [Seth Black](https://www.blackseoanalyzer.com/en/about) · Published September 17, 2026

Every few weeks someone asks me some version of “how do I see all the pages on a website?” The honest answer is that no single place has the full list. The sitemap has the pages the owner wants indexed. Google has the pages it chose to index. A crawler has the pages that are linked. Your CMS has the pages that exist in its database. Those four lists overlap, but they’re never identical, and the differences are usually where the problems are.

Here are the seven methods I use, starting with the fastest.

## The short answer

- **Check the sitemap:** open `example.com/sitemap.xml` (or look for the `Sitemap:` line in `example.com/robots.txt`).
- **Search Google:** type `site:example.com` to see pages Google has indexed.
- **Use Google Search Console:** the Pages report lists indexed and non-indexed URLs for sites you own.
- **Crawl the site:** a crawler follows every link from the home page and lists every page it reaches.
- **Look in your CMS:** WordPress, Shopify and others list every page and post in the admin.
- **Check analytics:** GA4’s Pages and screens report shows every page that got a visit.
- **Use the Wayback Machine:** its CDX API lists URLs it has archived for a domain, including old ones.

If you want one list you can trust, crawl the site and compare it with the sitemap. Details below.

## 1. Read the XML sitemap

Most sites publish a sitemap, and it’s the quickest way to see their pages. Try these URLs:

- `https://example.com/sitemap.xml`
- `https://example.com/sitemap_index.xml` (Yoast and Rank Math on WordPress)
- `https://example.com/wp-sitemap.xml` (WordPress core)
- `https://example.com/robots.txt`, then follow the `Sitemap:` line

Large sites use a sitemap index, which is a sitemap of sitemaps. Open each child file to get the URLs.

To pull the URLs out of a sitemap on the command line:

```bash
curl -s https://example.com/sitemap.xml | grep -o '<loc>[^<]*' | sed 's/<loc>//'
```

**The catch:** a sitemap is a list of pages the owner *wants* found. It’s often generated by a plugin that skips some post types, and it’s often stale. Pages that exist but aren’t in it are exactly the ones you need to know about.

## 2. Search Google with site:

Type this into Google:

```text
site:example.com
```

You can narrow it to a section with `site:example.com/blog` or a subdomain with `site:shop.example.com`.

**The catch:** `site:` shows a sample of indexed pages, not a complete list, and the result count is a rough estimate. It also only shows what Google indexed. Pages Google crawled and decided not to index won’t appear. That makes it useful for a quick look at someone else’s site, and not reliable for an audit.

## 3. Use Google Search Console (sites you own)

If you have access to the site in Search Console, go to **Indexing → Pages**. You’ll see URLs grouped by status: indexed, and not indexed with a reason (crawled but not indexed, duplicate without canonical, excluded by noindex, 404, and so on). Click a reason to see example URLs and export them.

Two limits to know about. The report shows example URLs, not necessarily every URL, and the export is capped at 1,000 rows per report. For large sites, add a separate Search Console property for each big section (`https://example.com/blog/`) to get a 1,000-row sample of each.

Bing Webmaster Tools has a similar Site Explorer view, and it’s worth checking too, because Bing sometimes finds URLs Google hasn’t.

## 4. Crawl the site

This is the method that finds what’s actually linked, which is what users and search engines can reach. A crawler starts at one URL, follows every internal link, and records each page with its status code.

With [bseoa](https://www.blackseoanalyzer.com/en/features), the crawler I build, this crawls a whole site and writes one line per page:

```bash
black-seo-analyzer \
  --url-to-begin-crawl https://example.com \
  --disable-external-links \
  --output-type jsonl-summary \
  --output-file pages.jsonl
```

Each line has the page’s `url`, `status_code` and warning count. To get a plain list of URLs:

```bash
jq -r '.url' pages.jsonl | sort > crawled.txt
```

If the site renders its navigation with JavaScript (React, Vue, Angular, a lot of Shopify themes), add `--spa`. Without rendering, a crawler sees the empty HTML shell and finds almost nothing. I wrote about that failure in [why your JavaScript crawl is broken](https://www.blackseoanalyzer.com/en/blog/why-your-javascript-crawl-is-broken-and-how-to-fix-it).

If you’d rather see the structure than a list, `--output-type sitemap` writes an interactive HTML map of the pages and how they link to each other.

Other options: Screaming Frog’s free version crawls up to 500 URLs, and `wget` can do a rough crawl if it’s all you have:

```bash
wget --spider -r -l inf -nd -nv -o crawl.log https://example.com
grep -o 'https\?://[^ ]*' crawl.log | sort -u
```

`wget` won’t render JavaScript and ignores most of what makes a page a page, but it will get you a list.

**The catch:** a crawler only finds pages that are linked from somewhere it can reach. Orphan pages, meaning pages that exist but have no internal links pointing at them, are invisible to it. That’s what the next step is for.

## Compare the crawl with the sitemap to find missing pages

This is the step most people skip, and it’s the most useful one. Save the sitemap’s URLs to a file, then compare them with the crawl:

```bash
# URLs listed in the sitemap
curl -s https://example.com/sitemap.xml | grep -o '<loc>[^<]*' | sed 's/<loc>//' | sort > in-sitemap.txt

# in the sitemap but not reachable by links: orphan pages
comm -23 in-sitemap.txt crawled.txt

# reachable by links but missing from the sitemap
comm -13 in-sitemap.txt crawled.txt
```

If the site uses a sitemap index, run the `curl` line on each child sitemap and append to the same file before sorting.

The first list is pages you told Google about but don’t link to, which usually means they rank badly or are leftovers that should be removed. The second is pages people can reach that your sitemap generator forgot. While you’re there, check the status codes in `pages.jsonl`. Internal links pointing at 301s and 404s are a common find, and sitemaps that list redirected or deleted URLs are even more common.

There are more places URLs hide than sitemaps and links, like canonical tags, hreflang annotations, pagination and JavaScript-only routes. I went through all of them in [why your site audit tool is missing half your pages](https://www.blackseoanalyzer.com/en/blog/why-your-site-audit-tool-is-missing-half-your-pages).

## 5. Look in the CMS

If you run the site, the CMS knows about pages no crawler can find, including drafts, private pages and pages nobody linked.

- **WordPress:** Pages → All Pages and Posts → All Posts. Custom post types (products, case studies) have their own menus. Tools → Export gives you everything as XML.
- **Shopify:** Online Store → Pages, plus Products, Collections and Blog posts, which are separate lists.
- **Webflow, Squarespace, Wix:** the Pages panel lists static pages. CMS collection items (blog posts, products) are listed separately.

**The catch:** the CMS doesn’t know about URLs created by the server or plugins, like tag archives, paginated URLs, search results pages or parameter variations. Those are often the pages that cause duplicate content.

## 6. Check your analytics

In GA4, go to **Reports → Engagement → Pages and screens** and set the date range to a year or more. The default “Page path and screen class” dimension lists every path that had at least one visit. Switch it to “Page path + query string and screen class” to see parameter variations too.

This catches old landing pages, campaign pages and URLs with parameters that nothing else lists. Anything with traffic that isn’t in your crawl is worth investigating, since visitors are reaching it somehow.

## 7. Use the Wayback Machine for old or removed pages

The Internet Archive’s CDX API lists every URL it has captured for a domain, including pages that have since been deleted:

```text
https://web.archive.org/cdx/search/cdx?url=example.com/*&output=txt&fl=original&collapse=urlkey
```

This works on any site, not just yours. It’s particularly useful after a migration, when you need to know which old URLs still need redirects.

## Which method should you use?

| Method | Works on any site | Finds unlinked pages | Complete |
| --- | --- | --- | --- |
| sitemap.xml | Yes | Sometimes | No, only what the owner listed |
| `site:` search | Yes | Yes, if indexed | No, a sample |
| Search Console | Only yours | Yes | Close, but exports are capped |
| Crawler | Yes | No | Yes, for linked pages |
| CMS | Only yours | Yes | Yes, for CMS content |
| Analytics | Only yours | Yes, if visited | No |
| Wayback Machine | Yes | Yes, if archived | No |

For someone else’s site: sitemap, then a crawl, then the Wayback Machine. For your own site: crawl it, compare against the sitemap, and check the gaps against Search Console.

## Seeing all the links on a single page

Sometimes the question is smaller: what links are on *this* page? Open the page, press F12, go to the Console tab, and run:

```js
[...new Set([...document.links].map(a => a.href))].join('\n')
```

That lists every unique link after JavaScript has run. To get the links along with anchor text, internal or external status and the rest of the page audit, crawl just that one URL:

```bash
black-seo-analyzer \
  --url-to-begin-crawl https://example.com/page \
  --max-pages 1 \
  --output-type json \
  --output-file page.json
```

The links are in `link_analysis` inside each page’s `complete_analysis` block. If what you actually want is the dead ones, the [broken links report](https://www.blackseoanalyzer.com/en/blog/the-broken-links-report-that-actually-makes-sense) lists every broken link with the page it’s on.

## The part nobody likes to hear

Some pages can’t be found by any of these methods: pages with no links, not in any sitemap, not indexed, never visited, and never archived. If you’re auditing your own site, the CMS or the server’s file system is the only complete source. If it’s someone else’s site, those pages are effectively private, which is usually what the owner intended.

If you want to try the crawl-and-compare approach on your own site, [bseoa has a 14-day free trial](https://www.blackseoanalyzer.com/en/free-trial) with every feature included. It runs on your machine, on Windows, macOS or Linux.

-Sethers

About the author

[Seth Black](https://www.blackseoanalyzer.com/en/about)

Seth is a software engineer and engineering leader who builds bseoa, a Rust-based technical SEO crawler with a GUI and CLI. He writes about the crawling, rendering and indexing problems he runs into on real sites.

[LinkedIn](https://www.linkedin.com/in/seth-black-tx/) · [GitHub](https://github.com/sethblack) · [YouTube](https://www.youtube.com/@SethBlack)

[Back to Blog](https://www.blackseoanalyzer.com/en/blog)

## Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

[Download Free Trial](https://www.blackseoanalyzer.com/en/free-trial)

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://www.blackseoanalyzer.com/#organization","name":"Fiscus Technology, LLC","url":"https://www.blackseoanalyzer.com/","logo":{"@type":"ImageObject","url":"https://www.blackseoanalyzer.com/static/images/black-seo-analyzer.png"},"sameAs":["https://sethserver.com/","https://github.com/sethblack"],"founder":{"@id":"https://www.blackseoanalyzer.com/#seth"}},{"@type":"WebSite","@id":"https://www.blackseoanalyzer.com/#website","name":"bseoa","url":"https://www.blackseoanalyzer.com/","publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"}},{"@type":"Person","@id":"https://www.blackseoanalyzer.com/#seth","name":"Seth Black","url":"https://www.blackseoanalyzer.com/about","jobTitle":"Creator of bseoa","sameAs":["https://www.linkedin.com/in/seth-black-tx/","https://github.com/sethblack","https://www.youtube.com/@SethBlack","https://sethserver.com/"]},{"@type":"BlogPosting","@id":"https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website#article","headline":"How to See All Pages on a Website","name":"How to See All Pages on a Website","description":"See every page on a website using the sitemap, a site: search, Search Console, a crawler, your CMS, analytics or the Wayback Machine, and find orphan pages.","url":"https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website","datePublished":"2026-09-17T02:49:24","dateModified":"2026-09-17T02:49:24","author":{"@id":"https://www.blackseoanalyzer.com/#seth"},"publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"},"image":{"@type":"ImageObject","url":"https://www.blackseoanalyzer.com/static/images/black-seo-analyzer-og-image.png","width":1200,"height":630},"mainEntityOfPage":{"@type":"WebPage","@id":"https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website"},"isPartOf":{"@id":"https://www.blackseoanalyzer.com/#website"},"inLanguage":"en-US","isAccessibleForFree":true,"wordCount":1696},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://www.blackseoanalyzer.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://www.blackseoanalyzer.com/en/blog"},{"@type":"ListItem","position":3,"name":"How to See All Pages on a Website","item":"https://www.blackseoanalyzer.com/en/blog/how-to-see-all-pages-on-a-website"}]}]}
```
