Can AI Crawlers See Your JavaScript Content?
By Seth Black · Published
A client-rendered site can look perfectly healthy in Google and be close to invisible to ChatGPT, Claude, and Perplexity. Same URL, same content, very different results. The difference is whether the bot runs your JavaScript, and most AI crawlers don’t.
If your product copy, pricing, or docs only show up after a framework hydrates, this post covers who can see it, how to check, and how to fix it.
The short version
- Googlebot renders JavaScript. Google’s AI features in Search sit on top of that same crawling and rendering.
- The major AI crawlers from OpenAI, Anthropic, and Perplexity fetch raw HTML and don’t execute JavaScript, based on the best public evidence available.
- Nobody at those companies has published a spec saying so. Their crawler docs describe what each bot is for, not how it parses pages. What we know comes from observing their traffic.
That last point matters, so let me be precise about the evidence.
What the evidence says
The most cited source is Vercel’s analysis with MERJ, published in December 2024 from traffic across Vercel’s network. Their conclusion: “none of the major AI crawlers currently render JavaScript.” The list covered OpenAI’s GPTBot, OAI-SearchBot, and ChatGPT-User, Anthropic’s ClaudeBot, Meta’s crawler, ByteDance’s Bytespider, and PerplexityBot.
The interesting detail is that some of these bots do download JavaScript files. Vercel saw them in about 11.5% of ChatGPT’s crawler requests and about 23.8% of Claude’s. They just never ran them. A bot that fetches main.js and treats it as text is not a bot that sees your rendered page.
The same study found two exceptions. Gemini uses Googlebot’s infrastructure, which renders JavaScript. And Applebot renders JavaScript with a browser-based crawler, similar to Googlebot.
That study is getting old. Independent log analyses since then have reached the same conclusion, and I haven’t seen credible evidence that any of these crawlers started rendering. But these are observations, not guarantees, and any provider could change its pipeline tomorrow without announcing it. Build for the conservative assumption.
What about Google’s AI features?
Google’s guide to generative AI features in Search says those features “are rooted in our core Search ranking and quality systems,” and that “Google is able to process content within JavaScript as long as it isn’t blocked.” So AI Overviews and AI Mode work from what Google’s normal pipeline indexed, rendering included.
That doesn’t mean Google handles JavaScript perfectly. Rendering is deferred, resource-limited, and cached. I wrote about those gaps in what Googlebot actually renders versus what you think. But Google is at least trying to run your code. The AI crawlers above aren’t.
Agents are a different case
There’s a newer category that muddies this: AI agents that drive a real browser on a user’s behalf. OpenAI’s ChatGPT agent runs a cloud browser, and Google lists Google-Agent as a fetcher that navigates the web and takes actions when a user asks.
A browser-driven agent will see your rendered page. But those visits are one person’s task, not a crawl that builds the index an assistant cites from. For showing up in AI answers, the non-rendering crawlers are the ones that decide what exists.
What a non-rendering bot actually sees
Here’s a typical client-rendered page as it leaves the server:
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Loading...</title>
<script type="module" src="/assets/index-a1b2c3.js"></script>
</head>
<body>
<div id="root"></div>
</body>
</html>
A crawler that doesn’t execute JavaScript gets exactly that. A title of “Loading…”, an empty div, and a script tag it won’t run. No headings, no body copy, no internal links to follow, and no structured data if your JSON-LD is injected by a component.
When a user asks an assistant about your product, this page gives the assistant nothing to quote. At best it falls back to whatever other sites said about you.
How to test your own pages
You don’t need special tools to find this problem. You need to compare two versions of the same page.
View Source versus Elements
In Chrome, right-click and choose View Page Source. That’s the raw HTML from the server, which is what GPTBot and ClaudeBot get. Then open DevTools and look at the Elements panel. That’s the DOM after JavaScript ran.
If your H1 and main copy are in Elements but not in View Source, non-rendering bots can’t see them. I walk through this test in more detail in why your crawler can’t see your React site.
Disable JavaScript
In DevTools, open the Command Menu with Ctrl+Shift+P (Cmd+Shift+P on a Mac), type “Disable JavaScript,” and reload. What’s left on screen is roughly what a non-rendering bot gets. Remember to turn it back on.
curl with the bot’s user agent
This also tells you whether your server or CDN treats bots differently:
curl -s -A "Mozilla/5.0 (compatible; GPTBot/1.4; +https://openai.com/gptbot)" \
https://www.example.com/pricing | grep -i -c "per month"
Pick a phrase that should be on the page. A zero means the raw HTML doesn’t contain it. Check the status code too. A 403 means a firewall rule is blocking the bot before rendering even comes up.
Diff raw against rendered
For a real answer, save both versions and diff them. Headless Chrome can dump the rendered DOM:
curl -s https://www.example.com/pricing > raw.html
google-chrome --headless=new --dump-dom https://www.example.com/pricing > rendered.html
diff <(sed 's/></>\n</g' raw.html) <(sed 's/></>\n</g' rendered.html) | less
The binary might be chromium or chrome depending on your system. The sed just breaks tags onto separate lines so the diff is readable. --dump-dom prints the DOM once the page loads, so content that arrives after a slow API call can still be missing. For those pages, use a script that waits for a specific element.
Doing this for one URL is easy. Doing it for a whole site by hand isn’t. Black SEO Analyzer can crawl the site twice: once with its plain HTTP crawler, and once with --spa, which renders every page in headless Chrome. Compare the two crawls and you can see what a non-rendering bot gets versus what a browser gets across every URL instead of spot-checking five.
Google’s view
Search Console’s URL Inspection tool shows the rendered HTML Google produced. Read the HTML, not the screenshot. It only tells you about Google, though, which is the one pipeline that renders. A pass there says nothing about ChatGPT or Perplexity.
How to fix it
The goal is simple to state: everything you want an AI assistant to read and cite must be in the HTML response, before any JavaScript runs.
Server-side rendering or static generation
This is the real fix. Next.js, Nuxt, SvelteKit, Remix, Astro, and Angular’s SSR support all render HTML on the server, then hydrate for interactivity. Static generation is even better for pages that don’t change per request. Marketing pages, docs, and blog posts are all good candidates.
Google lists server-side rendering, static rendering, and hydration as the approaches to use. It calls dynamic rendering, serving bots a separate prerendered version, “a workaround and not a recommended solution.”
Prerender the pages that matter
If you can’t move the app to SSR right now, prerender the routes that carry your business: home, pricing, product and category pages, docs, comparison pages. Many build tools can output static HTML for a list of routes. Serve that same HTML to everyone. Serving different content to bots than to users is how you end up with a cloaking problem.
Put the critical content in the initial HTML
Even with SSR, sites leak content back to the client. Check each of these is in the raw HTML:
<title>, meta description, and canonical- The H1 and the main body copy
- Prices, specs, and availability on product pages
- Navigation and internal links as real
<a href="...">elements, not click handlers - JSON-LD structured data, rendered server-side instead of injected by a component. Schema markup is only useful if the bot can find it.
- FAQ answers and tab content present in the markup, not fetched when someone clicks
Interactive pieces like filters, carousels, and account widgets can hydrate later. Nobody cites your carousel.
Don’t count on <noscript> or embedded JSON
A <noscript> block with a summary is better than nothing, but it’s easy for it to drift away from the real page. Frameworks also embed state as JSON in a script tag, like Next.js’s __NEXT_DATA__. Whether any AI crawler parses that is undocumented. Don’t bet your visibility on it.
Check that robots.txt isn’t doing the blocking
If a page is server-rendered and still never shows up in AI answers, look at access before you look at rendering. A wildcard block or a WAF rule can stop the search crawlers entirely. The AI crawler robots.txt guide covers which tokens to allow. An llms.txt file won’t fix a rendering problem either, because anything it links to still has to return real HTML.
The takeaway
Assume every AI crawler other than Google’s and Apple’s reads only your raw HTML. Test with View Source, curl, and a raw-versus-rendered diff. Fix it with SSR, static generation, or prerendering of the pages that matter.
None of this is new advice. It’s the same thing that’s made JavaScript sites easier for Google to index for years, which is why SEO for AI keeps turning out to be plain SEO.
See what a non-rendering bot sees on every page, not just the five you remembered to check. Black SEO Analyzer runs locally with a CLI and GUI. Start a 14-day free trial with every feature, no page limit, and no credit card, or see one-time pricing.
-Sethers