What is a website crawler?
A website crawler (also called a web crawler, spider or bot) is a program that visits a web page, reads its HTML, collects the links on it, and then visits those links, repeating until it has found every page it can reach. Search engines use crawlers to discover pages for their index. Googlebot is the best-known example.
An SEO crawler, or SEO spider, does the same thing on a single site so you can see it the way a search engine does. Instead of building an index, it reports what it finds on every page: broken links, redirects, missing or duplicate titles, canonical and hreflang problems, noindex tags, slow or oversized resources, and so on. It's the starting point of almost every technical SEO audit.
What bseoa checks on every page
Each crawled page goes through 16 analysis modules. The main ones:
- Metadata: missing meta descriptions, missing Open Graph and Twitter Card data, and invalid or untyped JSON-LD. A
--serp-modecrawl adds title and description checks for missing, short and truncated snippets. - Links: links that return 4xx or 5xx errors, invalid URL formats, insecure HTTP links, empty, generic or URL-only anchor text, and tracking parameters in internal links.
- URLs: uppercase characters, unsafe characters, excessive length, depth and query parameters.
- Internationalization: missing or invalid
langattributes, duplicate hreflang values, invalid hreflang URLs and missing x-default. - Performance and Web Vitals: render-blocking scripts and stylesheets, oversized pages, scripts and images, missing image dimensions, lazy-loading and preload hints for LCP elements, font formats and
font-display, caching and compression headers. - Mobile: viewport problems, small tap targets, small font sizes, and responsive image issues.
- Security: mixed content, insecure form actions, third-party scripts without Subresource Integrity, and inline scripts and event handlers.
- JavaScript and CSS:
document.write,eval, deprecated APIs, synchronous scripts in the head,!importantoveruse and duplicate selectors. - Content: thin content, readability, keyword stuffing, and on-device semantic similarity between pages to spot cannibalization, with no page content sent to a third party.
Reference pages explaining the fix for each issue are in the technical SEO issue library.
Internal link checker
Because the crawler records every link on every page with its anchor text and whether it's internal, a crawl doubles as an internal link audit:
- Broken internal links:
--output-type broken-linksexports each dead link with the page it's on and its anchor text. See the broken links report. - Which pages link to a page: query the JSON export for every page that links to a URL. The commands are in how to find links to your website.
- Orphan pages: compare the crawl with your sitemap, analytics or logs to find pages nothing links to. See how to find orphan pages.
- Site structure:
--output-type sitemapwrites an interactive HTML map of how pages link to each other.
GUI or command line
The desktop GUI shows crawls as they run, keeps every crawl, and offers command suggestions as you type. The CLI does the same work from a terminal, a cron job or a CI pipeline:
black-seo-analyzer \
--url-to-begin-crawl https://example.com \
--spa \
--output-type html-folder \
--output-file ./audit
Exports include JSON, JSONL, a compact per-page JSONL summary, XML, CSV, flat CSV, HTML reports (with your own templates), a broken links CSV and an interactive site map. Every crawl is stored in SQLite, so you can export again later without re-crawling, list stored crawls, and compare two crawls to see what changed after a release. There's also a read-only MCP server for asking an AI assistant questions about crawl data. See the CLI documentation.
Crawl settings you control
- Concurrent requests and rate limiting, so you don't overload a server.
- A custom user agent string.
- A maximum page count, or no limit.
- Crawling from an XML sitemap instead of a start page.
- Skipping external link checks.
- Ignoring robots.txt when auditing sections of your own site that it disallows.
How bseoa compares with other SEO spiders
Desktop crawlers like Screaming Frog and Sitebulb and cloud audit tools like Semrush and Ahrefs all crawl sites. The main differences are where they run, how they're priced and what limits they put on crawls. bseoa is a one-time purchase that runs locally. For current prices and crawl limits across the market, see SEO crawler pricing compared. There are also head-to-head write-ups for Screaming Frog, Sitebulb, Semrush Site Audit and Ahrefs Site Audit.
Is there a free website crawler?
bseoa has a 14-day free trial with every feature and no credit card. For one-off checks of a single page, the free SEO tools cover robots.txt, XML sitemaps, schema markup, hreflang, canonical tags and title lengths without installing anything. Screaming Frog's free version crawls up to 500 URLs without JavaScript rendering.