How to Find and Fix Every Broken Canonical Tag on Your Site

By · Published · Updated

How to Find and Fix Every Broken Canonical Tag on Your Site

Canonical tags break quietly. No 500 errors, no alerts. Rankings slip on a handful of pages over a few weeks, everyone blames the algorithm, and when somebody finally opens the source, the canonical points at a URL that 301s to a URL that canonicals back to the first one. Google stopped trying to make sense of it a month ago.

A broken canonical is worse than a missing one. With no canonical, Google picks something reasonable on its own. With a bad one, you’re actively feeding it wrong information and then acting surprised when it listens.

Why canonicals break in weird ways

A canonical is one line of HTML, and it lives in templates. Templates get touched by everyone who has ever had CMS access: the engineer who ran the HTTPS migration, the marketer who added a tracking parameter, the contractor who set up pagination six years ago and now works somewhere else. Nobody owns the canonical rule. It just sits there collecting damage.

A few ways I’ve watched it happen:

  • An HTTP to HTTPS migration ships, and the canonical template keeps pointing at http:// for another year.
  • A campaign parameter handler rewrites canonicals to include ?utm_source=.
  • A staging or CDN hostname sneaks into the template during a redesign.
  • A canonical gets injected by a tag manager or a client-side hook, so the raw HTML says one thing and the rendered page says another.

The page looks fine in a browser. View Source looks fine if you don’t read carefully. Google quietly decides your product page isn’t the one worth indexing.

The failure modes that cover almost everything

Canonical to a redirect. The target 301s somewhere else. You told Google “the real version is over there,” and over there isn’t a real page.

Canonical to a non-indexable URL. The target returns 404 or 5xx, carries a noindex, or is blocked by robots.txt. You pointed at a ghost.

Chains. Page A canonicals to B, and B canonicals to C. Google doesn’t follow canonicals like redirects. It picks one and moves on, and it’s rarely the one you wanted.

Loops. A chain that ate itself. A canonicals to B, B canonicals back to A. Or A to B to C to A. Every hop returns 200 and every tag is valid HTML. I had a client whose entire blog category system sat in a three-node loop for over a year with organic traffic to those URLs close to zero. We went through CMS settings on two calls before anyone traced the canonical past the first hop.

Cross-domain canonicals that shouldn’t be. Usually staging.yoursite.com or a CDN hostname. Occasionally intentional, usually a template bug.

Self-reference problems. The page should canonical to itself and doesn’t, or it points at the trailing-slash variant, which 301s back to the non-slash version.

Some real patterns I’ve hit more than once:

  • Category page canonicals to a filtered version, which canonicals back to the category page.
  • A product canonicals to a variant URL that 404s because the variant was discontinued.
  • Paginated pages all canonical to page 1, and page 1 canonicals to /shop/ because someone thought that was cleaner.

Why spreadsheets stop working

Crawl, export, sort. That’s fine on a 500-page site. Somewhere around 10,000 URLs it falls apart, because you’re not looking for bad canonicals, you’re looking for bad chains. That means resolving the target, checking its status code, checking its canonical, checking that target, and repeating until you land on a 200 that canonicals to itself. By the third VLOOKUP you’ve lost the thread.

This is a graph problem. Every URL is a node, and every canonical is an edge. A clean canonical is either a self-loop or a single edge to a node that self-loops. Anything longer is a bug, and a cycle is a worse bug.

Mapping the canonical graph

Crawl with JavaScript rendering on if any part of the site is client-rendered. If a canonical is injected after load, the raw HTML won’t show you the one Google eventually sees.

I use Black SEO Analyzer for this. It flags malformed canonicals on every page (empty values, query strings, fragments, non-HTTPS targets) and stores each page’s final HTML, status code, and redirect target in a local SQLite file. Resolving the chains is then about twenty lines of Python:

black-seo-analyzer --url-to-begin-crawl https://yoursite.com --spa --db-path canon.db
import re, sqlite3
from urllib.parse import urljoin

db = sqlite3.connect("canon.db")
pages = {url: (status, redirect, html) for url, status, redirect, html in
         db.execute("SELECT url, status_code, redirected_to, html_content FROM pages")}

def canonical_of(url, html):
    for tag in re.findall(r"<link\b[^>]*>", html or "", re.I):
        if re.search(r"rel=[\"']?canonical", tag, re.I):
            m = re.search(r"href=[\"']([^\"']+)", tag, re.I)
            return urljoin(url, m.group(1)) if m else None

for url, (status, redirect, html) in pages.items():
    if status != 200:
        continue
    path, target = [url], canonical_of(url, html)
    while target and target != path[-1]:
        if target in path:
            print("LOOP", " -> ".join(path + [target])); break
        if target not in pages:
            print("NOT CRAWLED", url, "->", target); break
        t_status, t_redirect, t_html = pages[target]
        if t_status != 200 or t_redirect:
            print("BAD TARGET", t_status, url, "->", target); break
        path.append(target)
        target = canonical_of(target, t_html)
    if len(path) > 2:
        print("CHAIN", " -> ".join(path))

Sort the output by type, then by URL pattern. Most sites I audit have somewhere between a few dozen and a few thousand problems, and they cluster in two or three templates.

What to fix first

Group by template before you fix anything. Canonical bugs almost never live in one URL. They live in /products/*, /blog/category/*, or /search?*. A bad rule in the product template breaks 40,000 URLs at once. Fix the rule and a thousand URLs heal in one deploy.

Within that, start with the templates that get the most internal links or the most organic traffic. A broken canonical on a page nobody visits is a chore. A broken canonical on your top product pages is costing money right now.

While you’re in there, run the broken links report too. Broken links and broken canonicals tend to pile up in the same neglected templates.

What a clean canonical graph looks like

  • Every indexable page returns 200 and self-references.
  • No canonical points at a redirect, a 404, a noindex page, or a URL blocked by robots.txt.
  • Chain length is 1. Never 2.
  • No cross-domain canonicals unless you genuinely meant it.
  • Paginated pages self-canonical. Filtered URLs either self-canonical or stay out of the index.
  • The canonical in the raw HTML matches the canonical after rendering.

Keep it from coming back

Don’t trust the deploy. Re-crawl the affected templates and confirm the canonical is what you think it is in the rendered HTML. Then treat it like a test. Run the same script against staging before a release and fail the build if a new chain, loop, or bad target shows up. A programmer reviewing a PR can see “this change introduced 312 canonicals pointing at redirects” before it ships, instead of six weeks later when traffic is already down.

Canonicals are boring, which is exactly why nobody watches them until something hurts. Fix the template, put the check in your pipeline, and go work on something more interesting.

Run the canonical audit on your own site. BSA’s free 14-day trial includes every feature with no page limit and no credit card.

-Sethers

Discover hundreds of SEO Issues in Seconds

Without Monthly Subscriptions

Comprehensive technical SEO analysis powered by ML and 16 specialized modules. Optional AI-powered insights from Claude, GPT-4, or Gemini. Get actionable insights in seconds, and never pay monthly fees again.

Download Free Trial

I use AI to generate images for my posts and for general editing, updates, and ironically SEO purposes. I used to draw all of the images for my personal blog (taleas) myself, but as the volume of content I produce has increased, I've turned to AI tools to help create visuals that complement my writing. I go out of my way to generate images that look strange, and don't represent real people. If you ever want to chat about my use of AI, please reach out.