---
title: "Ecommerce Index Bloat Audit: Find and Remove Junk URLs"
description: "Count URLs three ways, find indexed pages that never earned a click, and pick one signal per pattern. The index bloat audit I run on ecommerce sites."
image: "https://cdn.blackseoanalyzer.com/blog-images/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat.jpg"
canonical: "https://www.blackseoanalyzer.com/en/blog/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat"
language: "en"
---

# How to Audit an Ecommerce Site for Crawl Waste and Index Bloat

By [Seth Black](https://www.blackseoanalyzer.com/en/about) · Published September 16, 2026

![How to Audit an Ecommerce Site for Crawl Waste and Index Bloat](https://cdn.blackseoanalyzer.com/blog-images/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat.jpg)

Most ecommerce index bloat isn’t hiding. It’s loud and expensive, and it’s sitting in a log file nobody pulled and a Search Console export nobody sorted. Thousands of URLs Google has indexed that have never earned a click, competing with the category and product pages that pay the bills.

Here’s the audit I run when a store’s rankings are sliding and nobody has looked at what’s actually in the index.

## Step 1: Count URLs three different ways

Pull the product and category export from the CMS. Pull the XML sitemaps. Pull 30 days of verified Googlebot requests from the access logs. Compare the three numbers.

If the CMS says 14,000 products and Googlebot requested 310,000 distinct URLs, the audit is basically over. Something is generating crawlable URLs and nobody told the SEO team. Usually it’s a faceted filter, a sort parameter, or a session ID from a checkout experiment that shipped years ago and never got cleaned up.

The gap between those three numbers tells you most of what’s wrong before a crawler touches the site.

## Step 2: Group the waste by pattern

A spreadsheet of 300,000 URLs sorted alphabetically isn’t an audit. You want the top 10 to 20 URL patterns by request count, grouped by parameter or path segment:

- `?sort=*` - 48,000 requests
- `?filter.v.price.*` - 91,000 requests
- `/collections/all?*` - 22,000 requests
- `?utm_source=*` - 11,000 requests on product pages

Now you can make decisions per pattern instead of per URL. A few lines of Python over the log file get you this table. If facets dominate the list, the [faceted navigation audit](https://www.blackseoanalyzer.com/en/blog/auditing-faceted-navigation-find-and-fix-your-crawl-budget-sinkhole) walks through sorting every parameter into keep, canonical, block, or wait.

## Step 3: Find indexed pages that have never earned a click

Export Search Console’s performance report by page for the last 90 days. Put it next to the list of indexed pages from the Page indexing report. Any URL that’s indexed with zero impressions for 90 days is thin, duplicated, or targeting a query nobody searches.

On one Shopify store, a large share of indexed product URLs had zero impressions over three months. They were color and size variants that should have been canonicaled to the parent product.

The usual suspects on ecommerce sites:

- **Variant URLs** without a canonical to the parent product.
- **Internal search results** (`/search?q=`) linked from a “popular searches” widget.
- **Tag and collection archives** the platform generates automatically.
- **Paginated category pages** deep enough that nobody reaches them but Google.
- **Out-of-stock and discontinued products** left live with no policy.
- **Tracking and currency parameters** that escaped into internal links or canonicals.

## Step 4: Pick one signal per pattern

A URL blocked in robots.txt can still end up indexed if something links to it, and Google can’t see a `noindex` on a page it isn’t allowed to fetch. That’s how you end up with a pile of “Indexed, though blocked by robots.txt” entries in Search Console.

Pick one approach per pattern:

- If it should never be indexed and nothing external links to it, block it in robots.txt.
- If it’s already indexed or has external links, let Google crawl it and serve `noindex` until it drops out. Block it later if you want.
- If it’s a duplicate of a clean URL, canonical it and remove the internal links pointing at the parameter version.
- If the page is gone for good, return 410.

Mixed signals on the same pattern force Google to guess, and it usually guesses wrong.

## Step 5: Decide what pagination is for

Google stopped using `rel="next"` and `rel="prev"` as an indexing signal in 2019. If `/category?page=47` is in the index, Google is treating it as its own page.

Paginated pages list different products, so canonicaling all of them to page 1 is a hint Google may ignore. The cleaner options are to let them self-canonical and make sure products are reachable through shallower paths, or to keep deep pages crawlable but out of the index. Pick one. “Do nothing and hope” is how you got here.

## Step 6: Crawl and verify

Once the rules ship, run a full crawl and confirm the site now matches your policy: variants canonical to parents, search pages carry `noindex` or aren’t linked, no internal links point at blocked patterns. I use Black SEO Analyzer here because every page’s status code, redirect target, and HTML go into a local SQLite database, so checking “every URL matching this pattern has the canonical I expect” is a script instead of a sampling exercise. Then check the logs again in two weeks. Request counts on the junk patterns should fall, and crawl frequency on real products should rise.

## If you only have an hour

Do Step 1. If Googlebot is requesting more than three times the URLs your CMS knows about, stop everything else and close that gap. The rest can wait. For the full ecommerce audit order after that, see [technical SEO for ecommerce sites](https://www.blackseoanalyzer.com/en/blog/technical-seo-for-ecommerce-sites).

Want to crawl your whole catalog and check it against your index policy? [Try BSA free for 14 days](https://www.blackseoanalyzer.com/free-trial), every feature, no page limit, no credit card.

-Sethers

About the author

[Seth Black](https://www.blackseoanalyzer.com/en/about)

Seth is a software engineer and engineering leader who builds bseoa, a Rust-based technical SEO crawler with a GUI and CLI. He writes about the crawling, rendering and indexing problems he runs into on real sites.

[LinkedIn](https://www.linkedin.com/in/seth-black-tx/) · [GitHub](https://github.com/sethblack) · [YouTube](https://www.youtube.com/@SethBlack)

[Back to Blog](https://www.blackseoanalyzer.com/en/blog)

## Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

[Download Free Trial](https://www.blackseoanalyzer.com/en/free-trial)

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://www.blackseoanalyzer.com/#organization","name":"Fiscus Technology, LLC","url":"https://www.blackseoanalyzer.com/","logo":{"@type":"ImageObject","url":"https://www.blackseoanalyzer.com/static/images/black-seo-analyzer.png"},"sameAs":["https://sethserver.com/","https://github.com/sethblack"],"founder":{"@id":"https://www.blackseoanalyzer.com/#seth"}},{"@type":"WebSite","@id":"https://www.blackseoanalyzer.com/#website","name":"bseoa","url":"https://www.blackseoanalyzer.com/","publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"}},{"@type":"Person","@id":"https://www.blackseoanalyzer.com/#seth","name":"Seth Black","url":"https://www.blackseoanalyzer.com/about","jobTitle":"Creator of bseoa","sameAs":["https://www.linkedin.com/in/seth-black-tx/","https://github.com/sethblack","https://www.youtube.com/@SethBlack","https://sethserver.com/"]},{"@type":"BlogPosting","@id":"https://www.blackseoanalyzer.com/en/blog/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat#article","headline":"How to Audit an Ecommerce Site for Crawl Waste and Index Bloat","name":"How to Audit an Ecommerce Site for Crawl Waste and Index Bloat","description":"Count URLs three ways, find indexed pages that never earned a click, and pick one signal per pattern. The index bloat audit I run on ecommerce sites.","url":"https://www.blackseoanalyzer.com/en/blog/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat","datePublished":"2026-09-16T01:35:42","dateModified":"2026-09-16T01:35:42","author":{"@id":"https://www.blackseoanalyzer.com/#seth"},"publisher":{"@id":"https://www.blackseoanalyzer.com/#organization"},"image":{"@type":"ImageObject","url":"https://cdn.blackseoanalyzer.com/blog-images/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat.jpg","width":1200,"height":630},"mainEntityOfPage":{"@type":"WebPage","@id":"https://www.blackseoanalyzer.com/en/blog/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat"},"isPartOf":{"@id":"https://www.blackseoanalyzer.com/#website"},"inLanguage":"en-US","isAccessibleForFree":true,"wordCount":865},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://www.blackseoanalyzer.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://www.blackseoanalyzer.com/en/blog"},{"@type":"ListItem","position":3,"name":"How to Audit an Ecommerce Site for Crawl Waste and Index Bloat","item":"https://www.blackseoanalyzer.com/en/blog/how-to-audit-an-ecommerce-site-for-crawl-waste-and-index-bloat"}]}]}
```
