Do AI Crawlers Run JavaScript? What the Docs Actually Say

Vorec Team · 2026-09-27 · About 7 min read

"AI crawlers don't execute JavaScript" is one of the most repeated claims in AI-search advice, and it is used to justify expensive work: rebuild the docs site, add prerendering, move to server-side rendering, or the models will never see your content.

We went looking for a vendor document saying it. On the pages we checked, we did not find one.

This post is about what that does and does not tell you — and how to decide about rendering when the published evidence is thinner than the confident advice suggests.

Checked on 27 September 2026.

First, "AI crawler" is three different things

Most of the confusion comes from one phrase covering three categories with different economics:

CategoryWhat it doesExample user agents
Bulk/training crawlersFetch pages at scale to build datasetsGPTBot, ClaudeBot
Search crawlersBuild an index for an AI search productOAI-SearchBot, PerplexityBot
User-triggered fetchersFetch one page because someone asked right nowChatGPT-User

A browser-based agent driving a real browser session is a fourth case again, and it is the one most likely to render, because it is a browser.

There is no reason to assume these behave identically. A claim that covers all of them at once should be treated with suspicion — including the claim in this post's opening line.

What the crawler documentation actually covers

We read these pages on 27 September 2026:

They consistently document user-agent strings, how to allow or block each bot in `robots.txt`, and what each bot is for. None of the fetched text described whether the crawler runs a JavaScript engine.

What that supports: on those three pages, on that date, the question is not addressed.

What it does not support: that no vendor documents it anywhere, or that these crawlers do not render. Absence of documentation is not evidence of behaviour. If you find a vendor page that states it, that page beats this one.

A claim repeated in a hundred posts is not better evidenced than a claim repeated in one. Check whether any of them links to a vendor document, or whether they all link to each other.

What Google does document

Google documents its behaviour explicitly, which makes it worth separating from the rest.

Google's JavaScript SEO basics describes crawling, rendering and indexing as distinct stages, with rendering deferred until resources allow. Googlebot does execute JavaScript. Rendering happening in its own stage tells you it is treated as a separate step — it does not tell you what it costs, and it tells you nothing about any other company's crawler.

Conceptual illustration of requesting and rendering a page; it does not describe every crawler's behavior

Google also documents dynamic rendering — serving crawlers a prerendered version — as a workaround, not as recommended architecture. If you are choosing now, its guidance points toward server-side rendering, static generation or hydration rather than maintaining a separate crawler path.

The decision, with the conditions attached

It is tempting to present this as a decision that makes itself — server rendering wins whatever crawlers do, so stop worrying about the question. That is too neat. Here is the version with the conditions left in:

If crawlers render JSIf they don't
Server-rendered HTMLContent available immediately; still no guarantee of indexingContent available; still no guarantee of indexing
Client-side onlyMay work — but rendering can fail on timeouts, errors or blocked resourcesContent likely unavailable to that crawler

Two conditions the table needs:

HTML availability is not indexing. Serving content to a crawler is necessary, not sufficient. Pages can be fetched and still not indexed, for reasons that have nothing to do with rendering.

Server rendering has real costs. Infrastructure, caching complexity, cache-invalidation bugs, and a build or runtime path to maintain. "Just render on the server" is a genuine trade-off, not a free win.

What survives: client-side-only rendering has a failure mode that server-rendered HTML does not, and that failure mode is invisible to you unless you test for it. That is a reason to lean one way. It is not a proof.

What we do, and its limits

Our blog is a single-page React application with content in a database. We run middleware that detects crawler user agents and returns rendered HTML — article body, headings, tables and structured data — without requiring JavaScript.

Exactly what we tested, and when. On 27 September 2026, sending a Googlebot user-agent string with `curl` to `https://vorec.ai/blog/process-documentation-guide`, we received HTTP 200 with a rendered `<h1>` and one `Article` JSON-LD block present in the raw response body — no JavaScript executed. A must-differ control, `/blog/definitely-not-real-xyz`, returned HTTP 404 with no `Article` JSON-LD, confirming the check distinguishes a real article from a missing one rather than passing everything. A default user agent on the same real URL also returned 200.

The limits of that, stated plainly: it covers blog article routes, on that date, and it is a request we made ourselves. That is a request sample. A spoofed user agent is not an authenticated crawler, and what we receive is not proof of what any real bot receives. Confirming it requires more than a request you made yourself. Server logs show a user-agent string, which anyone can send, so attributing an entry to a genuine crawler means using whatever verification the vendor supports — published IP ranges or reverse-DNS where offered, and those resources differ between vendors. Even then, logs typically record status and bytes rather than the response body, so establishing what a crawler actually received means pairing verified requests with logged response evidence and, where available, the vendor's own rendered inspection tooling.

We are also not claiming every route on our site behaves this way. We tested blog article paths.

Two things we got wrong along the way, which are more useful than the parts that worked:

Returning 200 for pages that do not exist. The SPA served its app shell with a 200 status for unknown URLs, which looks like a real page. Blog article paths now return a genuine 404 with `noindex` when we can confirm the content is absent — and deliberately do not 404 when the lookup itself fails, because a transient database error must never tell a crawler that a live page was deleted. Concretely, the lookup returns one of three states — found, absent, or unknown — and only `absent` produces the 404; a failed fetch, an unparseable response or a thrown error all resolve to `unknown` and fall through to the previous behaviour.

Checking status codes instead of content types. A missing image also returned 200 with the HTML shell. `curl -I` looked healthy; the image was broken. Assert the `Content-Type`, not just the status.

On cloaking, since prerendering raises it: serving crawlers a rendered version of the same content users get is not generally treated as cloaking. The problem is serving materially different content. Keeping the two in parity is the thing to verify.

Inspecting a response body alongside a missing-page control

How to check your own site

That last step matters because a negative control can reveal a check that passes for the wrong reason.

FAQ

Should I block AI crawlers?

Training access and search access are separate decisions with separate user agents, and vendors document them separately. You can allow a search crawler while blocking a training crawler. Decide each deliberately rather than treating it as one switch.

Does Googlebot execute JavaScript?

Yes, and Google documents it. Googlebot is not the AI-specific crawlers; evidence about one is not evidence about the others.

Is dynamic rendering a good long-term answer?

Google describes it as a workaround rather than a recommended architecture. If you are building now, server-side rendering, static generation or hydration are the better-supported paths. We run a crawler path because of how our application is built, which is a constraint, not a recommendation.

How do I find out what is actually hitting my site?

Server logs filtered by user agent, plus whatever verification that vendor supports, for anything you intend to act on. Verification resources are not identical across vendors, so check what each one publishes.

Related reading: What Google Actually Says About llms.txt.

Want the recording to narrate itself? Record with Vorec — it has its own macOS recorder, and an AI agent can drive it for you — or upload a recording you already have. Vorec drafts narration matched to the workflow it captured and generates the voiceover, so nothing is spoken into a microphone. The same capture can also produce a written step-by-step guide. Start free — 7-day trial, 100 credits, no credit card required. Trial includes up to 3 projects; exports carry a watermark.

← Back to blog