How to Find Orphan Pages

By · Published

An orphan page is a page on your site that no other page links to. Visitors can’t click their way to it, and search engines can only find it through a sitemap, an external link or an old memory of the URL. Orphan pages tend to rank poorly, go stale, and pile up after redesigns and migrations.

The trick to finding them is simple once you see it: a crawler can’t find orphan pages, by definition. So you find them by comparing what a crawler can reach with lists of URLs that come from somewhere else.

Why orphan pages matter

  • They’re hard for Google to discover and value. Internal links are how search engines find pages and judge which ones matter. A page with zero internal links gets no link equity from your site.
  • Good content gets wasted. Old campaign pages, product pages removed from category listings, and blog posts dropped from pagination often still get searches.
  • Bad content lingers. Test pages, outdated prices and expired offers stay indexed long after everyone forgot them.
  • They signal a broken process. Orphans usually come from menu changes, CMS migrations, or templates that stopped listing some content type.

Not every unlinked URL is a problem. Landing pages for paid ads, thank-you pages and unsubscribe confirmations are often unlinked on purpose, and should usually be noindexed.

The method: crawl, then compare

  1. Crawl the site from the home page and save the list of URLs the crawler reached.
  2. Collect URL lists from other sources: the XML sitemap, Search Console, analytics, server logs and the CMS.
  3. Anything in the other lists but not in the crawl is a candidate orphan.

Step 1: crawl the site

Any crawler that follows links works. With bseoa:

black-seo-analyzer \
  --url-to-begin-crawl https://example.com \
  --disable-external-links \
  --output-type jsonl-summary \
  --output-file pages.jsonl

jq -r 'select(.status_code == 200) | .url' pages.jsonl | sort -u > crawled.txt

Add --spa if the site’s navigation is rendered with JavaScript. Otherwise the crawler misses links that users and Googlebot see, and you’ll get false orphans.

Normalize the lists before comparing. The same page can appear as https://example.com/page and https://example.com/page/, or with tracking parameters. Decide on one format and convert every list to it.

Step 2a: compare with the XML sitemap

curl -s https://example.com/sitemap.xml \
  | grep -o '<loc>[^<]*' | sed 's/<loc>//' | sort -u > sitemap.txt

comm -23 sitemap.txt crawled.txt > orphans-from-sitemap.txt

For a sitemap index, repeat the curl line for each child sitemap. Not sure where the sitemap is? The XML sitemap checker finds and validates it.

Step 2b: compare with Search Console

In Indexing → Pages, open “Indexed” (or “View data about indexed pages”) and export the list. Pages that Google has indexed but your crawl didn’t reach are orphans Google still knows about. Exports are capped at 1,000 rows, so for large sites add URL-prefix properties for individual sections to get more.

The Performance report is also worth exporting: pages that still get impressions but have no internal links are the first ones to re-link.

Step 2c: compare with analytics

In GA4, open Reports → Engagement → Landing page for the last 12 months and export it. Landing pages with sessions that aren’t in your crawl are being found through search, email or old links, but not through your site.

GA4 reports paths, not full URLs, so add your domain before comparing:

sed 's#^#https://example.com#' ga4-landing-pages.txt | sort -u > analytics.txt
comm -23 analytics.txt crawled.txt > orphans-from-analytics.txt

Step 2d: compare with server logs

Logs show every URL that bots and users actually request, including ones that appear nowhere else. On a combined log format, this lists paths Googlebot fetched successfully:

grep Googlebot access.log | awk '$9 == 200 {print $7}' | sort -u \
  | sed 's#^#https://example.com#' > googlebot.txt

comm -23 googlebot.txt crawled.txt > orphans-from-logs.txt

Verify Googlebot with a reverse DNS lookup if you’re filtering on the user agent alone, since anything can claim to be Googlebot.

Step 2e: compare with the CMS

Export every published page, post and product URL from the CMS (WordPress Tools → Export, a Shopify products CSV, or a database query). This is the only source that finds pages with no links, no traffic, no sitemap entry and no crawl history.

Step 3: merge the candidates

cat orphans-from-*.txt | sort | uniq -c | sort -rn > orphan-candidates.txt

URLs that appear in several sources are the most interesting: real pages that Google, users and your CMS all know about, but your navigation doesn’t.

What to do with each orphan page

Go through the list and put each URL into one of these groups:

Situation Action
Valuable page that should rank Add internal links from relevant pages, categories or navigation, and keep it in the sitemap.
Outdated duplicate of another page 301 redirect it to the current version.
Page that should exist but not be indexed (ads landing page, thank-you page) Add noindex and remove it from the sitemap.
Content that’s gone for good Return 404 or 410, remove it from the sitemap, and redirect only if an equivalent page exists.
Not really an orphan (linked from JavaScript the crawler didn’t render, or a URL format mismatch) Fix the comparison, not the site.

Then fix the cause. If orphans appeared after a menu change, update the menu. If a content type stopped appearing in category pages, fix the template. Otherwise the list grows back.

Preventing new orphans

  • Link new content on publish. Every new page should get at least one link from a related, already-indexed page, not just from the sitemap.
  • Recrawl after navigation changes and migrations, and repeat the comparison.
  • Keep sitemap generation and navigation in sync. A sitemap that lists pages the navigation doesn’t is an orphan report waiting to happen.
  • Watch deep pagination. Posts that only appear on page 40 of the blog archive are close to orphans. Category pages, related-post links and topic hubs keep them within a few clicks.

If you want to run the crawl side of this on your site, bseoa has a 14-day free trial with every feature. It runs locally on Windows, macOS and Linux.

-Sethers

Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

Download Free Trial

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.