Enterprise SEO Audit Strategies
By Seth Black · Published · Updated
Most “enterprise SEO” pitches I’ve sat through were about the logo on the invoice, not the work. Someone shows a dashboard with 40 widgets, promises an onboarding manager, and quotes you a number that assumes your legal team loves 90-day procurement cycles. Meanwhile the actual SEO team wants three things: crawl the whole site, parse the logs, and hand a list of fixes to a programmer before the next sprint.
Here’s how I think about audits at the scale where “just run Screaming Frog” stops being a serious answer.
Start with the logs, not the crawl
If you’re running a site with hundreds of millions of URLs, your crawler is guessing at what Googlebot cares about. Your logs are not guessing. They’re a receipt.
Before I run a single BSA crawl on a big site, I want log data. Which templates does Googlebot hit daily? Which sections haven’t been crawled in three months? Where is crawl budget getting burned on parameter soup? Petabyte-scale log parsing sounds intimidating, but the actual problem is storage cost per query and a parser that can verify bots. That’s it. Pick the right two tools and it stops being a science project.
Every large log analysis I’ve run lands on the same finding: Googlebot burns 60% of its time on pages nobody wants indexed, and the URLs you actually care about get hit once a quarter.
Crawl with rendering, or don’t bother
Large sites are almost always a pile of frameworks stacked on top of each other. A product page rendered in Next.js, a category page that’s still Rails, a help center on some SaaS subdomain, and a blog on WordPress because marketing wouldn’t give it up.
A crawler that only parses HTML is going to lie to you about half the site. With --spa, BSA renders each page in headless Chrome and captures what actually shipped. That matters when the engineering team changes hydration logic and forgets to mention it to SEO, which happens more often than anyone wants to admit.
For a billion-page site, you don’t crawl everything every week. You sample by template, crawl your priority sections in full, and diff each run against the last one. With BSA that means pointing the crawler at each priority section’s sitemap (--is-sitemap), capping long-tail batches with --max-pages, and comparing each run’s SQLite database against the previous one. On a recent billion-URL site, that meant full weekly coverage of about 4 million revenue-driving product pages, and rotating samples across the long tail so every template got seen at least once a month. You find the regressions that matter without paying to re-crawl a billion near-duplicate pages. If you’ve been evaluating tools that handle this kind of scale, the Screaming Frog vs Sitebulb desktop crawler ceiling is worth understanding before you commit to anything.
Dashboards the C-suite will actually open
Executives don’t want a 40-widget dashboard. They want three numbers and a trend line. Indexed pages that drive revenue, crawl coverage on priority templates, and the count of critical issues open longer than two weeks.
Everything else is a drill-down. Build the drill-downs for the SEO team, not for the board deck. BSA’s JSON and CSV output drops straight into whatever BI tool the company already has, so you can build that exact view without wiring together three separate tools to do it. If a VP asks “why is traffic down,” you don’t want to be clicking through 12 tabs to answer.
Scriptable output is not optional
At enterprise scale, nobody is logging into a tool to copy numbers into a spreadsheet. The crawler has to pipe findings into whatever BI or ticketing system the company already uses. Issues go into Jira, metrics go into Looker, and nobody is copying numbers into a spreadsheet by hand. That’s the whole point of scriptable output, whether it comes from an API or, as with BSA, from a CLI that writes JSON and CSV. Tools that lack this kind of scriptable output — the kind that OnCrawl struggles with at real dataset sizes — become a bottleneck rather than an asset.
Security review before anyone runs a crawl
For a hosted tool, SOC 2 isn’t a feature, it’s the cost of being allowed in the building. If your vendor can’t hand a security questionnaire back in a week, legal will kill the deal before the SEO team ever runs a crawl. BSA sidesteps most of that because it isn’t a hosted service: it runs on your own machines, and the crawl data stays there unless you turn on the optional AI insights with your own API key. There’s no vendor holding your URLs, so there’s far less for procurement to review. I got tired of watching good tools lose deals for reasons that had nothing to do with the tool.
Same loop, harder plumbing
The loop is the same whether the site has 200 pages or 2 billion: crawl, parse logs, find the problems and hand them to whoever can fix them, repeat. The only thing that changes at scale is how much the plumbing costs you when it breaks.
BSA ships a real CLI and JSON output. CLI docs — or grab the trial if you want it in CI.
-Sethers