Free Tool · No Signup · See What AI Sees

AI Crawlability Checker

See whether AI can access and read a page on your website.

Checks published rules for
GPTBot ClaudeBot PerplexityBot Google-Extended Applebot · and more

Usually takes 10–20 seconds. Public URLs only. Do not submit private, login-only, or secret links.

In plain words

What to do about it

Frequently asked questions

What is an AI crawlability checker?

An AI crawlability checker tests the technical signals that affect whether AI crawlers can access and read a web page. It can inspect crawler rules, server responses, raw HTML, indexing instructions, and content that depends on JavaScript.

BLURSOR reports each layer separately and shows the evidence behind its findings.

How do I check if a page is crawlable by AI?

Paste the page’s public URL into the checker and run the test. Review whether the relevant crawlers are allowed by robots.txt, whether the server accepts their user-agent strings, and whether the page contains useful text without JavaScript.

Test the exact page you care about. A crawlable homepage does not guarantee that every page on the site is equally accessible.

Is my website crawlable if Google can index it?

Not necessarily. Googlebot and AI crawlers can have different robots.txt rules, server treatment, and rendering capabilities.

A page that is accessible to Googlebot may still block GPTBot, ClaudeBot, PerplexityBot, or another crawler. The reverse can also be true.

Which AI crawlers does BLURSOR check?

We read the published robots.txt rules for crawler roles associated with OpenAI, Anthropic, Perplexity, Common Crawl, Google, Apple, Meta, Amazon, ByteDance, DuckDuckGo, Mistral, Microsoft, and others.

For live delivery, we use a representative set: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, Googlebot, Applebot, Meta-ExternalAgent, and Amazonbot.

Google-Extended is a control token, not a standalone crawler. User-directed, special-purpose, and other policy-only roles remain in the published-rules section without being presented as live probes.

Do AI crawlers run JavaScript?

Rendering capabilities vary. Some automated systems can process JavaScript, while others primarily use the HTML returned by the server.

BLURSOR reads the raw HTML and, when browser rendering is available, compares it with the finished page. A large difference can mean that some automated readers receive less content than a human visitor.

What can stop AI crawlers from reading a page?

Common blockers include crawler-specific robots.txt rules, firewall or CDN bot protection, 403 and 429 responses, authentication, redirect problems, accidental noindex instructions, and content that only appears after JavaScript runs.

The checker identifies which of these signals it can observe and separates confirmed findings from estimates.

Does passing the crawlability test mean ChatGPT will cite my page?

No. A passing result means BLURSOR did not find a technical access problem in the layers it tested. It does not prove that a vendor has crawled, indexed, cited, ranked, or recommended the page.

BLURSOR sends its probes from its own servers using published user-agent strings, not from OpenAI’s, Anthropic’s, Google’s, or Perplexity’s verified crawler IP ranges. A site that checks source IP addresses may treat a genuine crawler differently.

What the AI Crawlability Checker tests

AI crawlability is the ability of an automated system to reach a page and retrieve useful content from it. That depends on more than whether the page loads in your browser. Crawler-specific rules, server responses, indexing instructions, and JavaScript can all change what an automated reader receives.

BLURSOR checks those layers separately. A crawler blocked by robots.txt, a server returning 403, and a page whose content only appears after JavaScript are different problems. The report shows which one you have and the evidence behind it.

Crawler access rules

BLURSOR reads the site’s robots.txt and evaluates the rule that applies to each supported crawler.

These rules can differ by user agent. A site might allow an AI search crawler while blocking a training crawler, or allow Googlebot while saying nothing about GPTBot or ClaudeBot. The report keeps those roles separate instead of treating every AI-related bot as the same thing.

Server responses for selected crawlers

A permissive robots.txt does not guarantee access. A server, CDN, or firewall can still refuse a request based on its user agent, IP address, rate, or other signals.

We request the page with a representative set of published crawler user-agent strings and record each response. Other roles are evaluated only against robots.txt, and the report keeps those evidence types separate.

Content without JavaScript

A finished page in a browser can contain much more text than the HTML initially sent by the server.

BLURSOR extracts the readable text, headings, metadata, links, and structured data available in the raw HTML. When browser rendering is available, it compares that version with the fully rendered page so you can see what depends on JavaScript.

Indexing and control signals

The crawlability test also checks signals that can limit how a page is used after it has been fetched. These include robots meta tags, X-Robots-Tag headers, noindex, AI preference signals, and canonical URLs.

Access and indexing are separate. A crawler may be able to retrieve a page that the page itself asks not to be indexed.

How to make a page crawlable by AI

Review your robots.txt policy

If a crawler is disallowed, decide whether that rule is intentional. Search crawlers, user-directed fetchers, and training crawlers serve different purposes, so allowing one does not mean you must allow all of them.

Check your firewall and bot protection

If the page works in a browser but returns an error to a crawler probe, review your CDN, firewall, rate limits, and managed bot settings. These controls can override what robots.txt appears to allow.

Put important content in the initial HTML

If most of the page disappears without JavaScript, move the essential title, description, headings, and body copy into server-rendered or statically generated HTML. A crawler should not need to run your application before it can understand the page.

Remove accidental indexing blocks

Review noindex, X-Robots-Tag, canonical URLs, and crawler-specific directives. These signals are useful when intentional and costly when left behind by mistake.

What comes after crawlability

Being readable is only the first step.

BLURSOR checks whether AI crawlers can reach a page and whether its content is structured clearly enough to use. Those are diagnostic signals—not a promise that an AI system will cite or recommend it.

Beamtrace shows whether your brand appears in AI answers and provides in-depth recommendations for improving its visibility.

Explore Beamtrace →