SEO Audit CLI for Developers
By Seth Black · Published · Updated
Most SEO audits happen in a browser tab someone forgets to close. You log in, start a crawl, come back tomorrow, export a CSV, and email it to someone whose inbox is where CSVs go to die. Fine for a one-off. Useless if you want the audit to run on every deploy, or if the site you need to check isn’t on the public internet yet.
I’ve gotten the “can you audit our staging environment?” request more times than I can count. With most hosted tools the answer is some version of no, or yes if you open a hole in the firewall and hope. A command-line audit solves that and a pile of other problems you don’t notice until you try to automate something.
What a CLI audit gives you
Structured output. That’s the point. A GUI hands you a rendered report. A CLI hands you a file you can parse, diff, alert on, or load into a database. It also runs wherever you run it: your laptop, a build agent, or a box inside the VPN next to staging.
Here’s the basic shape of a run with Black SEO Analyzer:
black-seo-analyzer \
--url-to-begin-crawl https://staging.internal.example.com \
--max-pages 2000 \
--output-type json \
--output-file audit-$(date +%Y%m%d).json
No login, no render queue in another timezone. The process exits, and you have a file. Add --spa if the site renders content with JavaScript, so pages get rendered in headless Chrome before they’re analyzed. Other output types include jsonl, csv, xml, html-folder for a browsable report, and broken-links for a list of broken URLs with the pages that link to them.
Every crawl also lands in a local SQLite database (crawl.db by default). You can regenerate any output format later with --generate-output without crawling again, and querying crawls with plain SQL is often faster than any export.
Working with the output
The JSON is a pages array. Each page has its URL, status code, and a block per analysis module with its warnings. That makes the obvious questions one-liners:
# Every URL that returned 4xx or 5xx
jq -r '.pages[] | select(.status_code >= 400) | "\(.status_code) \(.url)"' audit.json
For large sites or CI, I prefer the compact JSONL summary: one line per page with the URL, status code, warning count, warnings, and redirect target.
black-seo-analyzer \
--url-to-begin-crawl https://staging.example.com \
--output-type jsonl-summary \
--aggregate-warnings \
--min-severity high \
--output-file audit.jsonl
# Which warnings show up most
jq -r '.warnings[]?.key' audit.jsonl | sort | uniq -c | sort -rn | head -20
# Everything that redirects, and where to
jq -r 'select(.redirected_to != null) | "\(.url) -> \(.redirected_to)"' audit.jsonl
--aggregate-warnings puts a site-level summary on the first line: every warning key, its count, and the affected URLs. --min-severity drops low-severity noise before it reaches your scripts.
Piping audits into CI
The useful part is treating canonical tags, robots directives, and broken internal links the way you treat a failing unit test. If a change orphans 400 product pages, the build should fail before production does.
The pattern is small:
black-seo-analyzer --url-to-begin-crawl "$STAGING_URL" \
--output-type json --output-file audit.json
python scripts/check_regressions.py audit.json baseline.json || exit 1
baseline.json is the last known-good crawl, committed to the repo. When something changes on purpose, you update the baseline in the same PR, and reviewers see the delta. The trap is regenerating the baseline on every run, because then you’re comparing the site to itself and will never catch anything.
Your regression script decides what counts as a failure, which means the rules live in code and are actually enforceable. I start with four, because these are the ones that cost traffic when they break:
- Indexability regressions. Pages that were indexable in the baseline and now carry
noindex. Staging flags leak into production constantly. - Canonical integrity. Canonicals pointing at redirects, 404s, or hostnames you don’t own.
- Internal link changes. Pages that lost all inbound internal links between releases.
- Status code drift. 200s that became 301s or 404s without anyone filing a ticket.
A fifth rule has saved me more than once: fail if the indexable page count drops by more than a couple of percent. A template change that orphans a thousand pages trips it before anyone opens Search Console.
If your team won’t accept a hard fail yet, run the check as a warning for two weeks. People stop arguing once it catches a real one. The full list of what to check before a merge is in the pre-deployment SEO checklist.
npm scripts for local checks
In a Node project, put it in package.json so nobody has to remember the flags:
{
"scripts": {
"audit:staging": "black-seo-analyzer --url-to-begin-crawl https://staging.example.com --max-pages 500 --output-type json --output-file audit.json",
"audit:report": "node scripts/parse-audit.js audit.json"
}
}
npm run audit:staging && npm run audit:report becomes muscle memory. The parse script does whatever you need: a summary table, a GitHub issue, a diff against the last run.
Alerts people actually read
Every CSV emailed to a marketing inbox is a CSV nobody reads. A webhook that fires when staging has new errors gets read in thirty seconds:
COUNT=$(jq '[.pages[] | select(.status_code >= 400)] | length' audit.json)
if [ "$COUNT" -gt 0 ]; then
curl -s -X POST -H 'Content-type: application/json' \
--data "{\"text\":\"Staging SEO audit: $COUNT URLs returning 4xx/5xx\"}" \
"$SLACK_WEBHOOK"
fi
Crude, but it works. Once you have JSON, loading key numbers into Postgres or your metrics system is maybe 20 lines of Python. I keep a rolling table per site, so when something breaks I can see exactly when it started instead of staring at a red number with no history.
Why local matters
Hosted audit tools come with crawl limits, rate limits, and sometimes a sales call between you and the answer. Running locally means you can audit a 50,000-URL site twice in a row without watching a usage meter, and staging behind a VPN is no problem because the crawler runs where you run it. For the code-level bugs worth catching this way, like minifiers eating meta tags or router-driven title swaps, see technical SEO for developers.
Run it, parse it, alert on it. Nobody has to remember to check a dashboard nobody checks.
BSA ships a real CLI with JSON, CSV, XML, and HTML output. Read the CLI docs, or try it free for 14 days with every feature and no page limit.
-Sethers