Measurement glossary
Plain explanations of the AI Crawlability Checker’s measurements, with examples, practical next steps, and the limits of each check.
Back to the checker25 guides
Alphabetical · Written for business decisions
No matching measurements. Try a shorter term or choose All topics.
- Browser usability
Accessible names
An accessible name is the label a browser exposes for a control, such as a link, button, or field. It helps people using assistive technology identify what the control does.
Also: control names · unnamed controls · button labels
- Citation signals
Answer position
Answer position is BLURSOR's estimate of where the first heading followed by substantial body text occurs in the extracted content. It is a structural reading aid, not a judgment that the passage answers a customer's question or will be cited.
Also: heading-plus-body pair · heading-plus-answer · first answer · answer location · content blocks · early text
- Citation signals
Citation signals and enough content to assess
Citation signals are limited observations of page content and structure, such as quotations, numeric patterns, links, dates, and heading placement. BLURSOR first checks whether the raw page contains enough recognized content blocks to run those observations.
Also: citation signals · citeability · enough content to assess · too short to judge · thin content
- Citation signals
Content date
A content date identifies when a page says it was published, modified, or reviewed. BLURSOR reports the first usable date found through its supported detection order. A detected date is a declaration, not proof that the information is current or accurate.
Also: publication date · published date · last updated · modified date · freshness · datePublished · dateModified · time datetime
- Reading & structure
Content representation
Content representation describes how the same page's information appears in its initial response and browser snapshot. BLURSOR's focused check looks for zero-valued figures replaced after rendering and substantial repeated text that largely disappears. It does not compare every sentence for agreement.
Also: raw placeholders · zero-value placeholders · animated counters · raw duplication · duplicate blocks · representation quality
- Access & discovery
Content Signals and declared uses
Content Signals are machine-readable declarations in robots.txt that describe a publisher's intended uses of content, such as search, AI input, or model training. They are distinct from standard path-level Allow and Disallow rules.
Also: Content-Signal · Content Signals · AI usage preferences · search ai-input ai-train
- Citation signals
External links
External links point from the checked page to a different website hostname. BLURSOR counts distinct qualifying web addresses inside selected content blocks. A detected link is not proof that its destination supports a claim.
Also: outbound links · outside source links · external source URLs
- Browser usability
Form labels
A form label identifies the purpose of an input, selection, or text area. A useful label is understandable to the visitor and associated with the field so assistive technology can identify it.
Also: field labels · unlabeled fields · input labels
- Reading & structure
Headings and page structure
Headings label a page's topic and sections through HTML levels such as H1, H2, and H3. They help readers navigate the content. BLURSOR counts these elements before and after JavaScript; the counts do not judge the usefulness of the headings.
Also: H1 · H2 · H3 · heading hierarchy · section headings · heading structure
- Browser usability
Image alt text
Alt text is the text alternative supplied through an image's alt attribute. It should convey the image's purpose or information in context. A deliberately empty alt attribute can be appropriate for a decorative image.
Also: alternative text · missing alt text · image descriptions
- Reading & structure
Indexing instructions
Indexing instructions are page metadata or response headers that tell supporting search engines how a page may appear in search. A noindex instruction requests exclusion from the index. It differs from robots.txt, which controls crawling rather than search inclusion.
Also: noindex · robots meta · meta robots · X-Robots-Tag · indexability · none directive · noai · noimageai
- Browser usability
Layout shift
Layout shift is movement of visible page content between rendered frames. Cumulative Layout Shift, or CLS, summarises unexpected movement. BLURSOR shows a bounded browser observation rather than a field measurement of the whole visitor experience.
Also: CLS · Cumulative Layout Shift · visual stability
- Access & discovery
llms.txt and optional reading guidance
llms.txt is a proposed Markdown file that introduces a website or section and points AI tools toward useful material. It is optional guidance, not an access-control rule or evidence that an AI service reads or cites the site.
Also: llms.txt · llms-full.txt · AI reading guidance · LLM documentation index
- Browser usability
Main-content landmark
A main-content landmark identifies the primary content of a page. HTML's main element or an appropriate main role can distinguish that area from repeated navigation, banners, and other surrounding interface content.
Also: main landmark · main content area · page landmarks
- Reading & structure
Meta description
A meta description is a short page summary stored in HTML metadata. Search engines may use it as a result snippet when it suits the query. BLURSOR checks whether the selected initial response contains one, not whether a search engine displays it.
Also: description tag · page description · search description · search snippet · SEO description
- Access & discovery
Missing-page responses and soft 404s
A missing-page response tells a client that a URL has no available page, usually with HTTP 404 or 410. A soft 404 can occur when an unavailable address instead returns a successful status or an unrelated normal-looking destination.
Also: soft 404 · 404 · 410 · catch-all page · unknown paths · missing-page responses
- Citation signals
Numeric signals
Numeric signals are text patterns that BLURSOR recognises as possible quantitative statements, such as percentages, prices, ranges, or proportions. They are detected formats, not verified statistics or a measure of research quality.
Also: numeric patterns · statistic signals · quantitative claims
- Access & discovery
Page delivery and request evidence
Page delivery describes what a particular request received: the requested page, a challenge, a refusal, a temporary response, or an error. BLURSOR observes its own request profiles; it does not verify visits from genuine vendor crawlers.
Also: page delivery · request profiles · crawler verification · security challenge · response differences
- Reading & structure
Page title
A page title is the document label provided by its HTML title element. Browsers use it in tabs, and search engines may use it when choosing a result's title link. It is separate from the main heading displayed inside the page.
Also: title tag · HTML title · document title · browser tab title · SEO title · title link
- Citation signals
Quotation signals
Quotation signals are formatting patterns that BLURSOR recognises as possible quoted passages. The count describes detected text and markup; it does not verify the speaker, attribution, accuracy, or credibility of a quotation.
Also: quote signals · quoted passages · blockquote
- Reading & structure
Readable text before and after JavaScript
Readable text is the text BLURSOR extracts from a page response after removing common code and markup. Comparing the initial response with a browser snapshot shows a difference in text volume, not how much important content is missing.
Also: raw text · browser text · rendered text · character count · text ratio · JavaScript content · server rendering
- Access & discovery
robots.txt and crawler rules
robots.txt is a public file that tells supporting automated crawlers which website paths they may request. Its rules express crawling preferences; they do not secure a page or prove that a crawler visited it.
Also: robots.txt · Robots Exclusion Protocol · published crawler rules · crawler policy · robots precedence
- Access & discovery
Sitemaps and page discovery
A sitemap is a published file that lists website URLs, or other sitemap files, to help supporting crawlers discover content. Finding a sitemap does not establish that a particular page was included, crawled, or indexed.
Also: sitemap · XML sitemap · sitemap index · sitemap.xml
- Reading & structure
Structured data before and after JavaScript
Structured data labels information about a page in a machine-readable format. BLURSOR looks for JSON-LD blocks in the initial response and browser snapshot. Detecting a block does not confirm that its contents are valid, accurate, or eligible for a search feature.
Also: JSON-LD · schema · schema markup · schema.org · machine-readable data
- Browser usability
WebMCP
WebMCP is an emerging browser interface through which a web application can expose described, structured actions to an AI agent. Detecting a reference or API object does not show that any action was successfully performed.
Also: WebMCP tools · browser agent tools · modelContext
A measurement is a starting point
These checks show what BLURSOR could observe about one page. They do not predict whether an AI service will cite or recommend your business. Each guide separates the result from the decisions it can support.