Fixing Duplicate Meta Descriptions at Scale

By · Published

Fixing Duplicate Meta Descriptions at Scale

Duplicate meta descriptions look solved about five minutes into an audit. You dedupe the exact matches, the report goes green, and meanwhile 40,000 product pages share a description that differs by two words. String equality catches the byte-for-byte copies. The template underneath is where the actual mess lives.

The exact-match duplicates

Identical strings on multiple URLs. Every crawler catches these. They usually come from a CMS default, an empty field falling back to a global description, or an editor who duplicated a page and forgot to change the meta.

Fix them and move on. This part isn’t interesting.

The near-duplicates most tools won’t flag

Here’s the one that slips through. A product template outputs:

Buy {product_name} at {store_name}. Free shipping on orders over $50. Shop our selection today.

Every product passes a uniqueness check because {product_name} changes. But 90% of the string is identical on every page. Google often ignores descriptions like that and writes its own snippet, especially when the on-page content is also templated. You lose control of the one piece of SERP copy you’re supposed to own.

The same thing happens with:

  • Faceted URLs that inherit the parent category’s description
  • Paginated pages (?page=2, ?page=3) where the description never changes
  • Sort variants like ?sort=price-asc that create new indexable URLs with the same meta
  • Location pages where only the city name swaps inside a 40-word template
  • Blog archives and tag pages that fall back to the site-wide description

None of these fail a string equality check. They need similarity scoring.

Finding them at scale

Here’s the pass I run on any site with more than a few hundred pages.

  1. Crawl the site and extract every meta description with its URL. Render JavaScript if the site sets meta tags client-side, or you’ll be scoring the shell’s default description.
  2. Group URLs by structural pattern with a regex on the path: all /products/* together, all ?page= variants together, all location pages together.
  3. Inside each group, strip out the variable parts you already know about (product name, city, brand) and compare what’s left.
  4. Flag anything above roughly 85% similarity as a near-duplicate cluster.
  5. Trace each cluster back to the template, field fallback, or parameter that produced it.

I run the crawl in Black SEO Analyzer, which flags missing, too-short, and too-long descriptions on every page and stores each page’s HTML in a local SQLite database. The similarity step is a short Python script on top of that:

import re, sqlite3
from difflib import SequenceMatcher
from collections import defaultdict

db = sqlite3.connect("crawl.db")
groups = defaultdict(list)
for url, html in db.execute("SELECT url, html_content FROM pages WHERE status_code = 200"):
    m = re.search(r'<meta[^>]+name=["\']description["\'][^>]*content=["\']([^"\']*)', html or "", re.I)
    if m:
        pattern = re.sub(r"/[^/?]+$", "/*", url.split("?")[0])
        groups[pattern].append((url, m.group(1)))

for pattern, rows in groups.items():
    base = rows[0][1]
    similar = [u for u, d in rows[1:] if SequenceMatcher(None, base, d).ratio() > 0.85]
    if len(similar) > 10:
        print(f"{pattern}: {len(similar) + 1} near-duplicates, e.g. {base[:80]}")

It’s crude. Comparing everything to the first description in a group misses some clusters, and the regex assumes name comes before content. It still finds the three or four templates responsible for most of the noise in a few minutes, which is all you need to start.

Fix the template, not the pages

If 300 product descriptions are 95% identical, don’t hand-write 300 unique metas. That’s a week of work that gets overwritten the next time someone bulk-imports a feed.

Fix the template logic. Add variables that pull differentiating content: a key attribute, the material, the use case, the price range, the number of reviews. If the template can’t produce enough variation, the source data is thin. That’s a content problem, and the meta description is just showing you where it lives.

For faceted, sorted, and paginated URLs, rewriting descriptions is usually the wrong fix. Those URLs shouldn’t be indexable in the first place. Canonical them to the parent or keep them out of the crawl, and the meta problem disappears along with the URL problem. The faceted navigation audit covers which of those to use per parameter.

What to prioritize

Start with the templates that carry revenue or traffic: product, category, and location pages. A near-duplicate cluster on 40,000 product pages matters. A cluster on 200 tag archives can wait, or those archives can leave the index entirely.

The fix almost always lives in the template, the field fallback, or the indexability rules, not in the meta field itself. For the bigger picture on catalog sites, where duplicate descriptions usually show up alongside variant and pagination problems, see technical SEO for ecommerce sites.

Want every meta description on your site in one database? Try BSA free for 14 days, every feature, no page limit, no credit card.

-Sethers

Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

Download Free Trial

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.