Fix Crawl Errors Without the Headache
By Seth Black · Published · Updated
Crawl errors are boring right up until a category page stops ranking and nobody knows why. A 404 on a blog post from 2019 is noise. A 404 on the category page your homepage links to costs money. The answer to “why did traffic drop” is almost always something dull: a 404 where a 200 used to be, a redirect chain that grew two hops last month, a 500 nobody noticed because the page still loads in a browser.
Here’s the order I work through crawl errors, from the first export to a fix that stays fixed.
Get the whole picture, not the Search Console slice
Search Console tells you something is broken, eventually, with a vague bucket label and a chart that updates when Google feels like it. It samples, it lags, and on a large site the counts are approximate. It’s fine for noticing a problem. It won’t tell you which redirect broke this morning.
You want two data sources side by side:
- A fresh crawl. It tells you what’s reachable right now, with current status codes, redirect targets, and the page that links to each broken URL.
- Server logs. They tell you what Googlebot actually requested and what it got back. Pull at least 30 days and filter to verified Googlebot by reverse DNS, since the user-agent string alone is trivially spoofed.
Those two lists overlap less than you’d expect, and the gap is often the whole problem.
Bucket by failure mode before you touch anything
Group by response code first. Not by URL, not by template. The shape of the problem tells you where to look.
- A wall of 404s usually means somebody changed a URL pattern without writing redirects. Check the git log for route changes in the last 30 days.
- 5xx errors mean the server is unhappy, or your crawler hit it too hard. Slow the crawl down and run it again before you page anyone. If they only show up for bots, look for a WAF rule or rate limiter that caught Googlebot in a net meant for scrapers.
- Timeouts don’t have a status code, which makes them easy to miss. They’re often worse than a 500 because the crawler just gives up and moves on.
- Redirect chains longer than one hop waste requests and leak signals. They grow from decisions made months apart by people who never talked.
- Redirect loops usually come from two systems both trying to help. The CDN forces HTTPS, the app forces a trailing slash, and neither backs down. One loop I found only triggered for logged-out users, which is exactly the state Googlebot crawls in.
- Soft 404s return 200 with “nothing here” content: empty category pages, zero-result search pages, a not-found template that forgot to set a status code. Nothing looks broken, which is why they pile up.
- Blocked resources break rendering. If robots.txt blocks
/static/,/assets/, or/_next/, Google can load the HTML but not the JavaScript and CSS that build the page.
Once you know which bucket is biggest, stop staring at the spreadsheet and fix something real.
Group by pattern, not by URL
A 400,000-row error export isn’t a to-do list. It’s usually a dozen real problems wearing a lot of costumes. If 82,000 errors all match /products/*/reviews/legacy-*, that’s one ticket.
Sort each bucket by URL pattern with a regex or the first couple of path segments. Then, inside each pattern, sort by inbound internal links. A 404 on a page with 400 internal links pointing at it gets fixed today. A 404 on a 2017 post with one inlink can wait, or 404 in peace forever.
The logs add the second priority signal. URLs Googlebot keeps requesting that return 4xx or 5xx are burning crawl budget right now. So are old redirected URLs Googlebot keeps cycling through because your internal links were never updated, which is common and easy to fix once you see it.
Decide the disposition before you write code
For each pattern there are three options:
- 301 to a live equivalent when one exists and the intent matches. Category moved, product merged into a parent, post rewritten under a new slug.
- 410 when the page is gone and there’s no reasonable replacement. Don’t redirect dead pages to the homepage. Google treats that as a soft 404 and you’ve made the problem harder to see.
- Fix it when the URL should return 200 and something is actually broken: a bad template, a missing database record, a link built from the wrong field.
If you can’t pick one in about thirty seconds per group, you don’t understand the group yet. Look at what links to it and what it used to be.
Write the rule, not the list
If you’re adding redirects one URL at a time, you’ll be doing it forever. Find the pattern and write one rule:
location ~ ^/products/old-category/(.*)$ {
return 301 /shop/new-category/$1;
}
For chains, point the original URL straight at the final destination and remove the middle hops. Fixing the oldest hop usually collapses the rest.
Validate before you ship
Before a fix goes live, I run a small script against a sample of the affected URLs on staging:
import requests
for url in open("fixes.txt").read().split():
r = requests.head(url, allow_redirects=True, timeout=10)
print(r.status_code, len(r.history), url)
If len(r.history) is more than 1 on a URL that should resolve directly, you still have a chain. If the status code isn’t what the ticket promised, the fix isn’t done. Grep the output for anything that isn’t clean and hand that short list back to the programmer. It has saved me from shipping fixes that weren’t fixes more times than I’d like.
The handoff is where audits die
You have a CSV. The programmer has a ticket system. Asking someone to translate 400 rows into tickets by hand is how you guarantee nothing ships this sprint.
Hand them patterns, dispositions, and the source pages that contain the broken links, because that’s where the fix usually lives. I run the crawl in Black SEO Analyzer for this part: the broken links report lists every broken URL with the page that links to it, and every status code and redirect target lands in a local SQLite database, so grouping by pattern is a query. Redirect chains that end on a canonical tag are worth checking in the same pass, since broken canonical tags do the same quiet damage and cluster in the same templates.
Re-crawl and check the math
Ship, then re-crawl the affected patterns the same day. Check the logs again in three to seven days to see whether Googlebot moved on.
If you fixed a pattern that should have removed 12,000 errors and the count dropped by 800, something else is generating them. A second template, a stale sitemap, or internal links you didn’t find. The answer is in the crawl data, not in another meeting about the crawl data.
That’s the whole loop. Export, bucket, group, decide, fix, verify. No 40-slide audit deck. The goal is a lower error count after the next deploy than before it.
Want to run this loop on your own site? Try BSA free for 14 days, every feature, no page limit, no credit card.
-Sethers