Why Your JavaScript Crawl Is Broken (And How to Fix It)
By Seth Black · Published · Updated
JavaScript rendering is where site audits quietly lie to you. The tool says the page is fine. The page is not fine. I’ve blamed content and backlinks for ranking drops that turned out to be hydration failures, and I’ve watched a client lose about a third of their organic traffic over two months while their agency kept emailing clean audit PDFs. The crawler they paid for read the HTML shell and called it healthy.
Here’s the gap between what you ship, what Googlebot renders, and what your crawler sees, plus the failures I check first.
What Googlebot actually does with JavaScript
“Googlebot renders JavaScript” is true, and that one sentence has caused a lot of grief. Teams read it, assume Chrome parity, and ship a React app that gets indexed as a loading spinner.
Google’s renderer does run your JavaScript. It also works differently from the browser on your laptop:
- Rendering can lag the crawl. Google fetches the HTML first and renders later. If your content only exists after rendering, there’s a window where Google has the shell and nothing else.
- Resources can time out. Large bundles, slow API calls, and third-party scripts may not finish before Google takes its snapshot.
- No persistent state. No cookies or localStorage carried between page loads, and no service worker that “just worked” in testing.
- Cached renders go stale. Pushing a fix is live for users immediately and irrelevant to Google until it comes back.
Then there’s your crawler, which is usually doing even less. Most audit tools do one of three things: fetch HTML only, render with a short timeout or a page cap, or render with an engine that doesn’t fire the same events a real browser does. The footnote about rendering the first 500 URLs is usually on page four of the docs.
Test your own site in five minutes
Pick a page that should rank and doesn’t.
curl -s https://yoursite.com/some-page | grep -i "<title\|<h1\|canonical\|name=\"description\""
That’s the raw HTML the server sent. Now open the same URL in Chrome and look at the Elements panel, which shows the rendered DOM. Compare them.
If your title, H1, canonical, meta description, structured data, or body copy only appear in the rendered DOM, you have a client-rendered page, and every crawler that reads source is auditing an empty room. Google Search Console’s URL Inspection tool is useful here too. Look at the rendered HTML it shows you, not the screenshot. The screenshot will look fine while the DOM underneath is missing your H1.
The rendering failures I check first
Hydration mismatches that rewrite metadata
React and Angular ship server markup and then hydrate it on the client. When the server’s HTML and the client’s first render disagree, the framework can throw a console warning and replace the DOM. If the title or canonical was in the server HTML and a client component overwrites it with something different, you now have two versions of your SEO metadata. Which one gets indexed depends on when the snapshot happens.
Canonicals and meta tags that break on deep links
React Helmet, Next.js Head, and Angular’s Meta service all work, and they all have failure modes. The meta description gets stuck on whatever the first route set. The canonical is correct when you click through from the homepage and wrong when the URL is fetched cold, which is exactly how Google fetches it. Open Graph tags injected after mount are invisible to social platforms that don’t run JavaScript.
Links that aren’t links
Googlebot follows <a href>. It does not click <div onClick>. If your navigation is built on click handlers, your crawl depth is one page. The same goes for route changes that update the URL but not the title or canonical.
Soft 404s from the router
The router shows a “not found” component while the server returns 200 for every path. Google figures it out eventually and drops pages, and your crawler reports thousands of healthy URLs that don’t exist.
Client-side redirects
Redirects done with window.location or router.push inside an effect mean Google sees the original page first and then maybe follows the redirect. Use server-side 301s.
Blocked resources
If robots.txt disallows /static/, /_next/, or your API path, Google can fetch the HTML but not the code that builds the page. I’ve watched that exact rule wipe out a site’s indexation in one deploy. Check what Googlebot is allowed to fetch, not what your browser can fetch.
Main-thread blockers and slow data
I audited a site where a chat widget attached a listener that blocked the main thread for four seconds. The HTML was there, the DOM was there, and the browser refused to paint. In single-page apps, the Largest Contentful Paint element is usually a hero that appears after a data fetch. The shell renders fast and the content shows up two seconds later. The fix is almost always boring: server-render the hero, preload the image, move the fetch to the server.
Lazy content that never loads
Sections that load on scroll or intersection may never trigger for a crawler that doesn’t scroll. If it matters for ranking, it shouldn’t depend on a scroll event.
Put the important stuff in the initial HTML
The fix isn’t “rewrite everything in Next.js,” even though that’s the first reply in the Slack thread. Decide what has to be in the server response and what can wait for hydration.
What belongs in the initial HTML:
- Title, canonical, meta description, hreflang
- H1 and the first chunk of body content
- Primary navigation as real
<a href>links - Structured data for the page type
- Image
srcfor anything above the fold
What can hydrate later: filters, carousels, personalization, the cart drawer. Googlebot doesn’t care whether “Add to wishlist” works. It cares that the product name, price, and description are there. If you can’t move to server-side rendering this quarter, prerender the critical routes. It isn’t elegant, but it stops Google from indexing your loading state.
Checking the whole site, not one URL
The manual check tells you whether something is broken. A crawl tells you where. For one URL, curl and DevTools are enough. For 40,000, you need both views for every page.
Black SEO Analyzer renders pages in headless Chrome when you pass --spa, waits for network activity to settle, and audits the DOM that actually rendered, with no page cap. The trick I use is to crawl twice, once raw and once rendered, into separate databases:
black-seo-analyzer --url-to-begin-crawl https://yoursite.com --db-path raw.db
black-seo-analyzer --url-to-begin-crawl https://yoursite.com --spa --db-path rendered.db
Then diff the title, canonical, H1, and link count per URL between the two SQLite files. Every URL where the rendered version differs is a page that depends on Google finishing its render. Every URL that only shows up in the rendered crawl is a page a non-rendering crawler never found. Hand that list to the programmers who own those components, grouped by route. For large React and Angular apps, the enterprise SPA audit playbook covers routing traps and bundle weight in more depth.
If the tool can’t execute JavaScript, it’s auditing a skeleton. Render the page first, then audit it.
BSA renders JavaScript in headless Chrome, so React and Angular sites audit correctly. See the features or try it free for 14 days, no page limit, no credit card.
-Sethers