Technical SEO for Ecommerce Sites

By · Published · Updated

Technical SEO for Ecommerce Sites

Ecommerce sites break in ways marketing sites never will. Add a color swatch and one product becomes twelve URLs. Ship a new category filter and the crawler finds a combinatorial explosion nobody budgeted for. The product team celebrates the launch, and a month later somebody asks why the category pages stopped ranking.

Here’s the order I run an ecommerce technical audit in, the platform traps that catch people, and how to hand the results to programmers so something actually ships.

Start with a URL inventory, not a crawl

Before any crawler touches the site, I want to know how many URLs the store thinks it has. Pull three numbers:

  • The product and category count from the CMS or product feed.
  • The URL count in the XML sitemaps.
  • The number of distinct URLs Googlebot requested in the last 30 days of server logs.

The gap between those numbers tells you most of what’s wrong. On one Shopify store, the sitemap listed 8,400 products, and the logs showed Googlebot spending most of its requests on /collections/all?filter.v.price.* variants. Nobody had configured the filter URLs because Shopify “handles it.” Shopify does not handle it.

If you can’t get logs, use Search Console’s page-level performance export as a weaker proxy. When the crawled or indexed URL count is several times the product count, audit the index bloat before anything else. The rest of the audit can wait.

Sort by revenue before you sort by anything else

The fastest way to waste a week on an ecommerce audit is to treat every URL the same. A canonical bug on a product doing $80K a quarter is a fire. The same bug on something that sold twice since 2022 can wait.

Pull GA4 or order data, tag every product and category URL with its last-90-days revenue, and join that to your crawl export. Then sort issues by revenue instead of by URL count. Any crawler that exports CSV or JSON works for this. The join is the part that turns a flat list of 4,000 issues into a plan.

Faceted navigation is where most of the damage lives

Size, color, brand, price, rating, availability, sort order. Multiply them and you get URL counts that don’t fit in a spreadsheet. The questions that matter:

  • Which facet combinations get organic traffic?
  • Which ones have inbound links?
  • Which ones are pure crawl trap?

Keep the first two as real, indexable pages with their own titles and descriptions. A “men’s running shoes” page might be a money page. A “running shoes + red + size 9.5 + in stock + sort by price” page is noise. Canonical or block the noise, and stop linking to it internally.

The platform fingerprints are predictable. On Magento, layered navigation adds ?product_list_order= and ?product_list_dir= to every category. On Shopify, it’s the filter.* parameters. On WooCommerce, it’s usually ?orderby= plus whatever filter plugin got installed. The full log-file method for deciding between robots.txt, canonicals, and noindex per parameter is in the faceted navigation audit, because getting that choice wrong is how people wipe out their money pages along with the junk.

Duplicate and near-duplicate product pages

I’ve seen all of these on the same site:

  • The same product living under multiple category paths, like /mens/shoes/runner-x and /shoes/runner-x, with the canonical flipping depending on which navigation path the crawler took. Google indexed both, and traffic split across two URLs for one SKU. The template fix took a programmer about an hour. Finding it was the hard part.
  • Color and size variants exposed as separate URLs with no canonical to the parent. A product with five colors and four sizes becomes 20 near-duplicates.
  • Session IDs, currency switches like ?currency=USD and ?currency=usd, or tracking parameters bleeding into canonical tags.
  • Two hundred SKUs sharing a 90%-identical description because someone bulk-imported a manufacturer feed.

The last one is a thin content problem at scale, and generic “duplicate content” warnings won’t tell you which feed caused it. Compare the main content area across products in the same category, not the whole template. Templated meta descriptions have the same problem, and near-duplicate meta descriptions need similarity scoring, not string equality.

Pagination canonicals

Shopify collections paginate with ?page=2. WooCommerce uses /page/2/. Google stopped using rel="next" and rel="prev" as an indexing signal years ago, and plenty of themes still ship with those and nothing else.

What I look for: paginated URLs that canonical to page 1 while listing different products (Google may ignore that canonical), page 2 canonicaling to page 1 while page 3 canonicals to page 2, and filtered page 1 canonicaling to unfiltered page 1 only when a specific filter is active. Export every paginated URL with its canonical and sort. The pattern breaks jump out.

Also look for paginated orphans. Nothing on the site links to ?page=7 anymore, but it got linked six months ago and Google hasn’t forgotten. Diff the URLs in your logs against the URLs your crawl can reach. Anything Googlebot still requests that your own site no longer links to is usually where the real bloat lives.

Out-of-stock and discontinued products

Every store has an out-of-stock policy. Almost none follow it consistently. I’ve audited sites that 404 some out-of-stock items, 302 others to the category, leave some live with a “notify me” button, and noindex the rest, depending on which team added the product.

A reasonable default: keep temporarily out-of-stock pages live and indexable with accurate availability, 301 permanently discontinued products to the closest replacement or parent category, and 410 the ones with no sensible replacement. Pick a policy, write it down where people will find it, and audit against it.

Internal search and tag pages in the index

/search?q= pages and auto-generated tag pages get indexed by accident all the time, usually because a “popular searches” widget links to them. They’re thin, they multiply, and they compete with your real category pages. Keep them out of the index and remove the internal links that feed them.

Product schema that doesn’t match the page

Every tool gives Product schema a green checkmark if the JSON-LD parses. That doesn’t mean you’re getting rich results. Check that offers.price, offers.priceCurrency, and offers.availability hold real values that match the visible page.

Two recurring failures. Themes that hardcode availability: "InStock" regardless of stock. And prices rendered by JavaScript after a cart or region check, where a non-rendering crawler only sees the server placeholder and you ship 0.00 to Merchant Center. Render the page the way a browser does and compare the schema to what’s actually in the DOM. On WooCommerce, also check whether two plugins are both injecting Product schema and disagreeing.

Speed on the templates that make money

Product detail pages and listing pages fail Core Web Vitals in predictable ways: an unoptimized hero image, a JavaScript carousel, image tiles without dimensions, and a tag manager container nobody has reviewed as a whole. I wrote up how to fix Core Web Vitals across an ecommerce catalog separately.

Hand off templates, not URLs

A standard audit says “23,000 pages have duplicate content” and stops. That’s one row, not a fix. The actual work is clustering those 23,000 pages into the five or six patterns causing them: a facet template, a sort parameter, a tracking parameter the cart appends, a currency switcher, a bad canonical rule on the product template.

“Fix the pagination canonical in collection.liquid” is one ticket with a clear target. “Fix 340 product URLs” is a list of grievances.

When I run this in Black SEO Analyzer, it renders every page in headless Chrome, runs all 16 analysis modules, and writes the results to CSV, JSON, or a local SQLite database, so grouping issues by URL pattern and joining them to revenue is a query instead of an afternoon of pivot tables. The audit is done when a programmer has a specific template to change and a re-crawl confirms the count went down.

Want to run this audit on your store? Try BSA free for 14 days, every feature, no page limit, no credit card. After that it’s a one-time license, not a subscription.

-Sethers

Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

Download Free Trial

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.