The practical brief
The finding
Across four generative search systems evaluated in 2023, citation recall averaged 51.5% of verification-worthy statements and citation precision averaged 74.5% of citations. [c1] [c3] [c6] [c12] [c14]
Why it matters
A citation beside an AI-generated statement about your business does not establish that the linked page supports everything the answer says.
| Measure | Reported four-system average |
|---|---|
| Fluency | 4.48 out of 5 across evaluated queries; human answer-text ratings with citations hidden [c9] [c13] |
| Perceived utility | 4.50 out of 5 across evaluated queries; human answer-text ratings with citations hidden [c10] [c13] |
| Citation recall | 51.5% of verification-worthy statements fully supported; four-system average across evaluated queries [c3] [c11] |
| Citation precision | 74.5% of citations credited as supporting associated statements under the study’s precision rule; four-system average across evaluated queries [c3] [c12] |
What to try
BLURSOR’s practical interpretation
Our interpretation: review business-relevant answers for source support and relevance, separately from whether your business appears.
Study boundary
This historical audit measured citation support, not factual truth, current product performance, business visibility or commercial outcomes.
Check support before treating a mention as credible
If an AI answer mentions your business, its links are not an automatic credibility check. The practical decision is whether customers can verify the claims—not simply whether your website appears.
Across four systems, only 51.5% of generated verification-worthy statements were fully supported by citations on average. That does not mean the remaining statements were false; it means their cited evidence did not establish full support.
What the researchers tested
Researchers evaluated Bing Chat, NeevaAI, perplexity.ai and YouChat using their first complete, single-turn responses collected between late February and late March 2023. Each system received 1,450 queries across 12 distributions, mixing historical search queries, open-ended questions, generated debate questions and how-to queries.
Human annotators rated fluency and perceived utility with citations hidden. Separately, they assessed citation support in the context of the query and full answer. Citation recall measures how much verification-worthy content is fully supported; citation precision measures whether citations support their associated statements.
The precision rule also credits partially supporting citations when citations collectively fully support a statement and none supports it fully alone. These measures concern verifiability—tracing statements to evidence—not whether that evidence is true.
Helpful answer text still had support gaps
Average fluency was 4.48 and perceived utility was 4.50 on five-point scales. With citations hidden, these ratings describe answer text, not displayed citations’ credibility. Citation precision averaged 74.5%, alongside the 51.5% recall result.
Citation recall ranged from 11.1% for YouChat to 68.7% for perplexity.ai; precision ranged from 63.6% for YouChat to 89.5% for Bing Chat. These historical results describe different aspects of support, not a current product ranking.
Across the four systems, precision and perceived utility were inversely correlated (r =−0.96). This observational association does not show that better citations cause less useful answers.
Faithful source wording can still miss the question
The authors investigated copying and close paraphrasing by comparing generated statements with annotator-selected supporting passages. Both BLEU and BERTScore—text-similarity measures—were positively correlated with system-level citation precision (r = 0.80 for each).
Their proposed explanation is a hypothesis: copying can improve source support while reducing perceived usefulness when the copied material is irrelevant to the query or surrounding answer. Faithfulness to a webpage and relevance to a customer’s question are different checks.
The similarity analysis included only fully or partially supporting citations, used evidence selected by annotators, and took the maximum similarity when a statement had multiple citations. It therefore does not describe every statement or establish a causal effect on usefulness.
Separate discovery, support, relevance and truth
Our interpretation is to keep discovery and credibility separate when reviewing AI answers. The study did not test edits that increase business mentions, improve recommendations or generate sales. A practical review can distinguish four issues without claiming a proven visibility tactic:
- Presence: record whether your business appears.
- Support: check whether cited pages back every material part of the associated statement.
- Relevance: ask whether supported material actually answers the customer’s question.
- Truth: verify important facts independently; a supporting source can itself be wrong or outdated.
The source and evidence
Evaluating Verifiability in Generative Search Engines
Study limitations and disclosures
- The audit covers first, single-turn responses collected in February–March 2023, not current systems or extended conversations.
- Citation support is not factual accuracy. Sentence-level evaluation can also combine several independently checkable claims within one judgment.
- Fluency and perceived-utility ratings were collected with citations hidden. They assess answer text, not the helpfulness or credibility of displayed citations.
- Different abstention rates complicate system comparisons: NeevaAI abstained on 22.7% of 1,450 queries, while each other system abstained on fewer than 0.5%.
- Each query-response pair was ordinarily annotated once. On 250 randomly sampled pairs assessed for agreement, pairwise agreement ranged from 82.0% to 94.6%.
- The similarity analysis included only fully or partially supporting citations, relied on annotator-selected evidence, and selected maximum similarity for statements with multiple citations. The copying/paraphrasing explanation for lower perceived utility is a hypothesis, not a demonstrated causal effect.
- Amazon Web Services supplied Mechanical Turk credits. The work was supported in part by the AI2050 program at Schmidt Futures, Grant G-22-63429.
Evidence behind this briefing
[c1] The study evaluated four commercial systems using their first complete, single-turn responses collected in late February–late March 2023; these are historical system results, not measurements of current products.
Page 4, §3.1 Evaluated Generative Search Engines · Read the study
[c2] Each system was evaluated on 1,450 queries across 12 distributions: five distributions with 150 sampled queries each and seven NaturalQuestions distributions with 100 each. The sample mixed historical search queries, open-ended questions, generated debate questions, and how-to queries.
Page 12, Appendix A, Table 3 caption · Read the study
[c3] Averaged across the four systems, citation recall was 51.5% of generated verification-worthy statements and citation precision was 74.5% of citations. These measure source support, not factual accuracy. Precision also credits partially supporting citations under the collective-support rule in §2.4.
Page 7, §4.2 Citation Recall and Precision; metric definitions in §§2.3–2.4, pages 3–4 · Read the study
[c4] Annotators rated responses highly for fluency and perceived helpfulness: averages were 4.48 and 4.50 respectively on five-point agreement scales. These are perceived-utility judgments, not demonstrated accuracy or business outcomes.
Page 6, §4.1 Fluency and Perceived Utility; scale definition in §2.2, page 3 · Read the study
[c5] Across the evaluated systems, citation precision and perceived utility were inversely correlated (r =−0.96). Bing Chat had the highest precision but lowest perceived utility, whereas YouChat showed the reverse. This is an observational system-level association, not evidence that accurate citations cause lower utility.
Page 8, §4.3 Citation Precision is Inversely Related to Perceived Utility · Read the study
[c6] The authors explicitly distinguish verifiability from factuality: checking whether cited sources support a statement does not establish that the statement is true.
Page 10, Limitations · Read the study
[c7] Annotators judged whether all information in a statement was supported by its citations, interpreting the statement in the context of the query and full response. This is a citation-support evaluation, not an independent factual-accuracy test.
Page 17, Figure 10, Step 3 annotation guidelines · Read the study
[c8] On a random sample of 250 query-response pairs, pairwise annotation agreement ranged from 82.0% to 94.6%, and F1 against majority consensus ranged from 91.0 to 97.3. Table 4 gives minima equal to 82.0% and 91.0, despite surrounding prose describing agreement as greater than those thresholds.
Page 20, Appendix E, Table 4 · Read the study
[c9] Across all evaluated queries, average human fluency ratings on a five-point Likert scale were 4.40 for Bing Chat, 4.43 for NeevaAI, 4.51 for perplexity.ai and 4.59 for YouChat; the reported overall average was 4.48. These measure fluent presentation, not factual accuracy.
Page 21, Appendix F, Table 5 · Read the study
[c10] Across all evaluated queries, average human perceived-utility ratings on a five-point Likert scale were 4.34 for Bing Chat, 4.48 for NeevaAI, 4.56 for perplexity.ai and 4.62 for YouChat; the reported overall average was 4.50. These are perceived helpfulness ratings, not verified accuracy.
Page 22, Appendix F, Table 6 · Read the study
[c11] Across all evaluated queries, citation recall was 58.7% for Bing Chat, 67.6% for NeevaAI, 68.7% for perplexity.ai and 11.1% for YouChat; the reported average was 51.5%. The table caption interprets low recall as generated statements not being fully supported by citations, rather than necessarily being false.
Page 23, Appendix G, Table 7 · Read the study
[c12] Across all evaluated queries, citation precision was 89.5% for Bing Chat, 72.0% for NeevaAI, 72.7% for perplexity.ai and 63.6% for YouChat; the reported average was 74.5%. This measures whether citations support their associated statements, not how often a business is discovered or cited.
Page 24, Appendix G, Table 8 · Read the study
[c13] Annotators rated response fluency and perceived utility on a five-point Likert scale while citations were hidden. These ratings therefore assess the answer text, not the helpfulness or credibility of displayed citations; this qualification should accompany comparisons between perceived answer quality and citation support.
Page 13 (7013), Appendix C, Annotation Interface, first-step description · Read the study
[c14] The authors hypothesize—not demonstrate causally—that copying or closely paraphrasing cited webpages helps explain the inverse relationship between citation precision and perceived utility. Faithful support from a source does not ensure that the copied statement answers the query or fits the response.
Page 2 (7002), §1 Introduction; explanatory hypothesis developed in §§4.3–4.4 · Read the study
[c15] The similarity analysis was restricted to citations providing full or partial support. Annotators selected the minimal supporting sentences from cited webpages, where such sentences existed; this was not an analysis of every generated statement or citation.
Page 8 (7008), §4.4 Generative Search Engines Closely Paraphrase From Cited Webpages · Read the study
[c16] The authors measured BLEU and BERTScore similarity between generated statements and annotator-selected supporting evidence. For statements with multiple citations, they selected the maximum similarity to any associated citation’s evidence.
Pages 8–9 (7008–7009), §4.4, BLEU/BERTScore calculation prose · Read the study
[c17] Across the four evaluated systems—Bing Chat, NeevaAI, perplexity.ai, and YouChat—both BLEU and BERTScore similarity were positively correlated with average citation precision (r = 0.80 for each metric). The authors interpret this association as suggesting that higher precision may partly reflect copying or paraphrasing, rather than establishing a causal effect on utility.
Page 9 (7009), §4.4, prose following Table 2 · Read the study