The practical brief

The finding

In a 2023 benchmark, ChatGPT answering with five retrieved passages achieved citation recall of 73.6 on ASQA, compared with 26.7 when answering without retrieved evidence and attaching citations afterward, on a 0–100 scale. [c1] [c2] [c3] [c5] [c7]

Why it matters

A citation beside a business claim does not establish that the source supports it, and stronger citation support does not necessarily mean a more correct answer.

Answering with evidence versus adding citations afterward
MeasureFive retrieved passages during answeringAnswer first, attach citations afterward
Citation recall73.6 on the 0–100 citation-recall scale; ASQA26.7 on the 0–100 citation-recall scale; ASQA [c2] [c3]
Business interpretationQualitative interpretation: stronger evidence support in this benchmarkQualitative interpretation: visible citations can leave claims unsupported [c3]
ChatGPT results on ASQA. Citation recall measures statement support from cited passages, not independently verified truth. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: before publishing or relying on an AI answer, check whether its sources support the specific claims, separately from checking correctness and usefulness.

Study boundary

These benchmark experiments tested specific 2023 models, not current commercial search products, brand discoverability or marketing outcomes.

Check the claim, not just the citation

Before using an AI-generated answer in customer-facing material or a business decision, distinguish a visible citation from actual support. The practical risk is an answer that looks documented but makes statements its cited passages do not justify.

Gao and colleagues found that writing an answer first and attaching citations afterward produced relatively strong correctness but poor citation support. On ASQA, a question-answering dataset, ChatGPT using five retrieved passages scored 73.6 for citation recall, versus 26.7 for the answer-first, cite-later approach. That 46.9-point gap concerns evidence support—not brand visibility or customer trust.

Evidence [c2] [c3]

What the researchers actually tested

The researchers introduced ALCE, a benchmark for evaluating answers with citations. They selected 1,000 development examples from each of three datasets: ASQA, QAMPARI and ELI5. Supporting material came from Wikipedia or Sphere, a web corpus, split into 100-word passages rather than whole pages.

They assessed fluency, answer correctness and citation quality separately. Citation recall means the extent to which cited passages collectively support statements in an answer; it is not independent verification that those statements are true in the world.

An automated natural-language inference model judged whether passages supported claims. Human evaluation checked selected outputs, with 100 examples sampled from ASQA and ELI5; the paper does not clearly establish 100 examples per dataset. Against human labels, automated citation-recall classification achieved 85.1% accuracy and citation-precision classification achieved 77.6%.

Evidence [c1] [c2] [c9] [c10]

Better support is not the same as better answers

Selecting among four generated responses using automated citation recall improved human-rated citation support on evaluated ELI5 outputs. ChatGPT’s recall rose from 50.8 to 59.7, and precision from 52.4 to 60.6, on the reported 0–100 scale. These were citation-quality judgments, not factual-accuracy or audience-preference scores.

The distinction matters: in ASQA shortcut tests, simply copying retrieved text produced nearly perfect citation scores but substantially worse correctness and fluency. A well-supported passage can still fail to answer the question.

More evidence was not uniformly better either. Moving from five to twenty passages improved GPT-4’s ASQA citation recall from 68.5 to 73.0, while ChatGPT-16K’s fell from 76.2 to 73.7. These results apply to the tested model versions, not every present-day AI system.

Evidence [c4] [c5] [c7]

A practical review process, not a visibility tactic

Our interpretation is to keep evidence review separate from answer review. For a business commissioning AI-assisted content or evaluating a vendor, a useful test is whether the answer addresses the question, whether its claims are correct, and whether each cited source actually supports the associated wording.

  • Open citations for consequential claims rather than treating their presence as approval.
  • Check qualifications and partial support: a source may justify only part of a sentence.
  • Ask vendors to report answer correctness and citation support separately.

Evidence [c2] [c7] [c11]

What this study cannot tell a marketer

The study does not identify a tactic for getting your business cited, recommended or discovered. It tests prompting approaches on bounded datasets with specific 2023 model versions, and excludes harder reasoning and coding scenarios.

Its automated evaluator can mistakenly penalize citations that provide partial support. Reranking also used only one seeded ChatGPT run because of cost, limiting conclusions about repeatability. Treat these findings as reasons to inspect evidence—not as a current ranking of commercial AI products.

Evidence [c5] [c6] [c11] [c12]

The source and evidence

Enabling Large Language Models to Generate Text with Citations

Tianyu Gao, Howard Yen, Jiatong Yu, Danqi Chen · 2023

Study limitations and disclosures

  • The experiments concern specific 2023 model versions and benchmark datasets, not live brand visibility, recommendations, customer preference or marketing performance.
  • Automated evaluation is imperfect: citation scoring can misjudge partial support, fluency scoring is length-sensitive, and generated ELI5 reference claims may omit valid answers.
  • ChatGPT reranking used only one seeded run because each run required four generations. Human evaluation covered a limited sample of selected outputs from ASQA and ELI5.
  • The acknowledgments disclose an IBM PhD Fellowship, NSF CAREER award IIS-2239290, a Sloan Research Fellowship and Microsoft Azure credits. Surge AI supported human evaluation; annotators were paid an average of USD 20 per hour.

Evidence behind this briefing

[c1] ALCE samples 1,000 development examples each from ASQA, QAMPARI, and ELI5 to evaluate existing LLMs’ citation capabilities; it does not supply citation-supervised training data. Its corpora are Wikipedia for the first two datasets and Sphere for ELI5, segmented into 100-word passages rather than whole web pages.

Page 3, §2 Task Setup and Datasets; corpus details in Table 1 and §2 · Read the study

[c2] Citation recall measures whether cited passages collectively support a statement, not independently verified real-world truth or audience preference. TRUE, an NLI model, evaluates entailment; statement scores are averaged within each response. Evaluation allows at most three citations per statement, and QAMPARI treats each listed entity as a statement.

Page 4, §3.3 Citation Quality; page 2, §2 footnotes 4–5 · Read the study

[c3] In these benchmark experiments, generating answers without retrieved passages and attaching citations afterward produced relatively strong correctness but poor citation support. On ASQA, ChatGPT VANILLA with five passages scored 73.6 for citation recall versus 26.7 for CLOSED BOOK plus POST CITE: a 46.9-point difference on the reported 0–100 scale, not a demonstrated effect on brand discoverability.

Page 7, §5.1 Main Results; page 6, Table 4 · Read the study

[c4] Selecting among four sampled responses using automated citation recall improved human-rated citation support for ChatGPT on the evaluated ELI5 outputs: recall rose from 50.8 to 59.7 and precision from 52.4 to 60.6 on the reported 0–100 scale. These are citation-quality judgments, not preference or factual-accuracy scores. Automated reranking scores may be biased.

Page 9, Table 9; sampling method in page 6, §4.3; bias qualification in page 7, §5.1 · Read the study

[c5] More retrieved passages did not uniformly improve performance: GPT-4 benefited from longer context in these experiments, whereas ChatGPT-16K did not. On ASQA, GPT-4 with five versus twenty passages had exact-match recall of 41.3 versus 44.4 and citation recall of 68.5 versus 73.0; ChatGPT-16K had exact-match recall of 36.1 at both passage counts and citation recall of 76.2 versus 73.7. These results concern the tested 2023 model versions, not all current AI-answer systems.

Page 7, §5.1 Main Results; page 8, Table 7; model versions in page 6, §5 · Read the study

[c6] The evaluation has measurement limitations: MAUVE is length-sensitive and potentially unstable; ELI5’s generated reference claims may omit valid answers; NLI accuracy constrains citation scoring, particularly partial support. The datasets exclude harder reasoning and coding scenarios, and experiments emphasize prompting rather than updating model weights.

Page 10, Limitations · Read the study

[c7] In ASQA shortcut tests, copying retrieved text produced nearly perfect citation scores without comparable answer correctness or fluency. This cautions against treating citation support alone as answer quality.

Appendix D, Table 11, PDF page 15 (6479) · Read the study

[c8] On ASQA, ChatGPT RERANK had citation recall/precision scores of 84.8/81.6 versus VANILLA's 73.6/72.5, while EM recall was 40.2 versus 40.4. Parentheses report standard deviations, but RERANK's zeros reflect only one seeded run, not demonstrated absence of variability.

Table 19, ASQA full results, PDF page 18 (6482); run-count qualification in Appendix G.6 · Read the study

[c9] Against human annotation labels, automatic citation recall and precision classification achieved 85.1% and 77.6% accuracy. Detecting irrelevant citations had 75.6% recall and 66.1% precision, with false positives attributed to inability to recognize partial support. These are evaluator-validation results, not answer correctness rates.

Appendix G.5, PDF page 17 (6481) · Read the study

[c10] Human evaluation used Surge AI workers paid an average of USD 20 per hour and randomly sampled 100 examples from ASQA and ELI5 for selected ChatGPT VANILLA, ChatGPT RERANK and Vicuna-13B VANILLA outputs. The wording does not explicitly establish 100 examples per dataset. Utility was rated on a five-point helpfulness/informativeness agreement scale, separate from citation support.

Appendix F and F.1, PDF page 15 (6479) · Read the study

[c11] Automatic citation precision can falsely penalize a citation that supports only part of a statement. The authors report poor results from prompting ChatGPT to judge partial support and defer a better supervised evaluator to future work.

Appendix E, Citation Recall Discussion, PDF page 15 (6479) · Read the study

[c12] The main-results appendix reports three random seeds generally, but only one seeded ChatGPT RERANK run because each run repeats generation four times and further experiments would incur substantial costs.

Appendix G.6, PDF page 17 (6481) · Read the study