Auditing at Scale: Crawling 10 Million Pages Without Crashing
By Seth Black · Published
The first large crawl I ran at real scale got to about 4.2 million URLs before the hosted tool we were paying for ran out of memory and quit. Support suggested retrying off-peak. The site had close to 10 million URLs that mattered, and a chunk of revenue lived on templates the crawl never touched. We sent a status report anyway. It was not our finest work.
Crawling at this size is a different job. The SEO rules are the same. The plumbing around them is not: how URLs are queued, where results are written, how fast you can hit the origin without taking it down, and what happens when one template starts returning 500s halfway through.
The frontier is the product
People obsess over parsers and rule engines. On a 10-million-page site, the URL frontier, meaning the set of URLs seen, pending, and done, decides whether you finish. A simple crawler with an in-memory queue and an in-memory “seen” set falls over somewhere in the hundreds of thousands, depending on RAM. After that you’re swapping, crashing, or silently dropping URLs without knowing it.
What to look for in any crawler you trust at this size:
- Normalized URLs before they’re queued. Parameter order, trailing slashes, case, and fragments. Without normalization, a faceted category page turns into millions of “unique” URLs and the frontier never empties.
- Hard limits you control. A maximum page count per run, and patterns you can exclude, so a calendar widget that links to next month forever doesn’t eat the job.
- Failures recorded, not skipped. Timeouts and 5xx responses should land in the results with a status, so you can re-run just those.
Storage that doesn’t fight you
The other place tools fall apart is the write path. Every URL produces a status code, headers, the HTML, outgoing links, and the analysis results. At 10 million URLs, anything that buffers the whole crawl in memory until the end is going to hurt.
- Stream results to disk as you go. If the process dies at 60%, you should still have 60% of the data.
- Index what you’ll query. You’ll ask “every 404 on the product template” a hundred times. Put indexes on URL, status code, and crawl session.
- Keep crawling and reporting separate. Crawl once, then generate as many reports as you want from the stored data without hitting the site again.
This is how Black SEO Analyzer is built. It’s written in Rust, runs on your own hardware, and writes every page to a local SQLite database as it crawls, with indexes on the columns you’ll filter on. Reports in JSON, CSV, XML, or HTML come out of that database with --generate-output, so you can re-slice a finished crawl without re-crawling. I wrote about why SQLite holds up for crawl data if you’re skeptical.
One giant crawl is usually the wrong shape
Past a few million URLs, a single breadth-first crawl from the homepage is fragile and slow to act on. Chain smaller jobs instead:
- Crawl priority sections first. Most large sites split sitemaps by type, so point the crawler straight at a sitemap file (
--is-sitemap) for products, then categories, then content. - Crawl the long tail in capped batches with
--max-pages, into separate databases if you like. - Throttle to what the origin can handle with
--concurrent-requestsand--rate-limit. A crawl that triggers the WAF or pages the on-call engineer ends early. - Merge findings by URL pattern and send the issue list to your ticket system.
Each step is small enough that if it breaks, you re-run that step instead of starting over.
Find orphan clusters by diffing, not crawling
On a site this big, orphans aren’t single pages. They’re whole neighborhoods: a category tree that lost its parent link in a migration, product pages reachable only through a sitemap marketing forgot about. You won’t find them by following links from the homepage.
Diff three lists:
- URLs your crawl reached through internal links
- URLs in your XML sitemaps
- URLs Googlebot requested in the last 30 days of server logs
Anything in the sitemaps or logs that the link crawl never reached is your orphan list. Anything the crawl reached that neither Google nor your sitemaps know about is often crawl-trap junk. If you’re wondering how your current tool handles this, your site audit tool may be missing half your pages.
Log data is non-negotiable
Crawling 10 million URLs tells you what’s broken. Logs tell you what Googlebot actually visits, and at this scale the gap between those lists is where the real work hides. A million orphan pages don’t matter much if Googlebot never asks for them. Two hundred broken canonicals on a template Googlebot hits every hour matter a lot.
Join the crawl results to parsed log data by URL and filter the issue list to “URLs Googlebot requested in the last 30 days.” That single filter usually cuts the list to something a programmer can ship in a sprint.
If you can’t get logs, because legal said no or the CDN charges per gigabyte to export them, use 90 days of Search Console data weighted by impressions as a proxy for what Google cares about. It’s worse than logs and much better than guessing. The broader process for running audits on sites this size is in enterprise SEO audit strategies.
What to ask any crawler before a big job
- Does it write results to disk as it goes, or hold everything in memory until the end?
- Is there a page cap, or a price that changes with page count?
- Can I throttle it to protect the origin?
- Can I start from a sitemap file instead of the homepage?
- Can I get raw results out as structured data to join with logs, or only a dashboard?
If most of those answers are no, you aren’t buying a crawler for a large site. You’re buying a very expensive sample.
BSA runs locally with no page cap and stores every crawl in SQLite. Try it free for 14 days, every feature, no credit card.
-Sethers