The practical brief

The finding

Across ten software prompts tested on all four engines, individual engines captured an average of just 11.4%–42.6% of the combined exact-URL citation set. [c1] [c4] [c6] [c7]

Why it matters

A dashboard observing one engine can miss sources cited elsewhere; a page-quality score alone does not establish live citation probability.

How much of the combined citation set did one engine capture?
EngineMean observed source coverage
ChatGPT42.6% of the four-engine exact-URL union, averaged across 10 complete prompts [c4]
Google24.3% of the four-engine exact-URL union, averaged across 10 complete prompts [c4]
Perplexity23.6% of the four-engine exact-URL union, averaged across 10 complete prompts [c4]
Microsoft Copilot11.4% of the four-engine exact-URL union, averaged across 10 complete prompts [c4]
Mean per-prompt coverage of the observed four-engine exact-URL union across ten complete prompts on 6 June 2026. These are source-coverage measurements, not traffic shares. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: require engine-specific coverage labels and distinguish page relevance assessments from observed citations before using either to allocate marketing effort.

Study boundary

This small observational benchmark measured displayed citations, not hidden retrieval, answer accuracy, customer trust or sales.

Check what a visibility dashboard actually measures

Before using an “AI visibility” dashboard to prioritize content spending, check whether it measures live answers, evaluates pages offline, or does both. These measure different things. In this audit, one engine was a weak proxy for the sources cited across four engines—not evidence that every business needs four-engine monitoring.

Benjamin Tannenbaum examined an internal benchmark of 15 English-language commercial prompts about AI search visibility software, run in a United States locale against ChatGPT, Microsoft Copilot, Google and Perplexity. The 6 June 2026 snapshot contained 589 citation observations, representing 528 exact URL strings and 356 domains. Only ten prompts had observations from all four engines.

Evidence [c1] [c4]

The same prompts produced different cited pages

Across 73 same-prompt engine-pair comparisons, 84.9% shared no exact cited URL. Mean Jaccard similarity was 0.0079. Jaccard measures shared items divided by all distinct items in the two sets; zero means no overlap. Domain overlap was also low, so different URLs on the same websites did not fully explain the result.

Among the ten complete prompts, no pair of engines shared an exact URL within its respective first five recorded citation positions in any of 60 comparisons. Those positions may mean different things across product interfaces, so this is supplementary evidence rather than a directly comparable ranking measure.

For those ten prompts, ChatGPT captured an average of 42.6% of the combined four-engine URL set; Google captured 24.3%, Perplexity 23.6% and Copilot 11.4%. These percentages describe observed source coverage, not market share, traffic or the probability that customers will encounter a business.

Evidence [c2] [c3] [c4]

Page fit is not the same as being discovered

The paper separates exposure—whether a page becomes available to an answer process—from selection—whether that process cites it once available. An engine-free score can assess page quality or query-page fit, meaning relevance to a particular request. It cannot automatically establish the chance of a live citation.

Likewise, an experiment that supplies candidate pages to a model tests citation preference after exposure. It does not test whether a production engine would discover those pages organically. This audit observed final citations, so it cannot identify which hidden stage caused the differences.

Our practical interpretation is to use page assessments as diagnostics and engine observations as outcome measurements, rather than treating either as a substitute for the other.

  • Ask which engines, prompts, locale and collection dates a visibility metric covers.
  • Keep offline relevance scores separate from displayed-citation results.
  • Consider broader monitoring where additional engines matter to your audience; this study does not establish the optimal monitoring mix.

Evidence [c1] [c6] [c7]

What the study cannot tell your business

The benchmark is small and unusually self-referential: software for measuring AI visibility. Its overlap rates should not be transferred to other industries without replication. Exact URL matching can also miss equivalent or syndicated content. No pages were experimentally edited, so the audit does not establish that any content optimization tactic works.

A separate comparison found 67.0% mean URL-set turnover across 41 matched prompt-engine pairs from 5 to 6 June. That is a diagnostic snapshot, not an expected daily volatility rate: collection changes could contribute, Google lacked earlier observations, and duplicate 4–5 June snapshots were excluded.

Tannenbaum is affiliated with Aiso Boost Ltd. and is its founder. The company develops AI visibility software and supplied the collection infrastructure, creating a direct commercial conflict. Aggregate data and figure code are available, but the full operational log is not. The paper contains no separate funding statement. The findings do not establish citation accuracy, substantive answer influence, trust or conversions.

Evidence [c5] [c7] [c8]

The source and evidence

Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores

Benjamin Tannenbaum · 2026-09-19

Study limitations and disclosures

  • The benchmark covers 15 English-language AI visibility software prompts in a United States locale; only ten were observed on all four engines. Its uncertainty estimates describe this benchmark, not all industries or information needs.
  • Only displayed citations were observed. Hidden retrieval and selection stages were unavailable, so the audit cannot identify what caused the differences between engines. No pages were experimentally changed, and citation presence does not establish accuracy, trust, answer influence or downstream sales.
  • Exact URL matching can treat equivalent content as different sources. Citation positions may also have different meanings across interfaces, limiting rank-sensitive comparisons. Source-coverage percentages are not estimates of market share or user traffic.
  • The 5–6 June turnover comparison may reflect engine changes, collection changes or both, and is not a general daily-volatility estimate. Duplicate 4–5 June snapshots were excluded; Google lacked earlier observations.
  • Author Benjamin Tannenbaum is affiliated with and founded Aiso Boost Ltd., which develops AI visibility software. Its infrastructure collected the benchmark, creating a disclosed direct commercial conflict. The paper contains no separate funding statement.
  • Released aggregates and figure-generation code allow reproduction of the reported figures and descriptive quantities, but the full operational citation log is withheld. Independent readers therefore cannot audit raw citation extraction or URL canonicalization.

Evidence behind this briefing

[c1] The internal benchmark covers 15 English-language AI visibility software prompts in a United States locale. Its 6 June snapshot contains 589 citation observations, 528 exact URL strings and 356 domains; only ten prompts have all four engines. Engine-free page fit and observed citation visibility are distinct measurement targets.

Section 4.1, Benchmark and collection window; counts in Table 1 · Read the study

[c2] Across 73 same-prompt engine pairs on 6 June, mean exact-URL Jaccard similarity was 0.0079, with a prompt-clustered bootstrap 95% interval of 0.0037–0.0132. Most pairs shared no exact URL. Domain overlap was also low; these are citation-set similarities, not measures of answer accuracy or source quality.

Section 5.1, Cross-engine exact-URL overlap is nearly zero · Read the study

[c3] Prominent citations did not form a shared exact-URL core: among ten prompts present on all four engines, none of the 60 engine-pair comparisons shared a URL within the first five recorded citation positions. Rank-weighted similarity was lower than full-set similarity. Citation-position semantics may differ across interfaces.

Section 5.2, Agreement does not reappear at the top of the citation list; Appendix A, Table 5 · Read the study

[c4] Single-engine monitoring captured only 11.4%–42.6% of the observed four-engine URL union on average across ten complete prompts. The mean per-prompt union expansion relative to the broadest engine was 2.43 times. This supports labelling engine coverage explicitly, not treating one engine as representative of all AI search; the percentages are not market share or user traffic.

Section 5.3, A single engine is a weak proxy for the multi-engine source universe; Appendix A, Table 4 · Read the study

[c5] Across 41 matched prompt-engine pairs for ChatGPT, Copilot and Perplexity, mean exact-URL turnover from 5 to 6 June was 67.0% (95% prompt-clustered bootstrap interval 61.1%–72.2%). Same-engine adjacent-day similarity was about 42 times same-day cross-engine similarity. This is a diagnostic comparison, not a daily-volatility estimate or causal engine effect.

Section 5.8, Same-engine sets move over time, but engine identity dominates this comparison; Appendix B · Read the study

[c6] An engine-free score can assess page quality or query-page fit, but interpreting it as end-to-end citation probability requires an exposure model or evidence that exposure is sufficiently invariant. Fixed-candidate experiments estimate selection conditional on exposure, rather than organic discovery.

Section 2, A Multistage View of Generative Visibility · Read the study

[c7] Only displayed citations were observed, so divergence cannot be attributed causally to retrieval, reranking or another hidden stage. The study did not manipulate pages. Its small, self-referential software benchmark cannot establish population-wide overlap rates, and strict URL matching can miss equivalent content.

Section 8, Displayed citations are not retrieval traces; related qualifications under Narrow domain, Small prompt set, Exact URL identity is strict, and No causal content intervention · Read the study

[c8] The author discloses a direct commercial conflict: he founded Aiso Boost, whose infrastructure collected the benchmark. Aggregate data and figure code are released, but the full operational citation log is not, preventing independent auditing of raw extraction and URL canonicalization.

Section 8, Commercial conflict of interest; Data and Code Availability · Read the study