The practical brief
The finding
For citation lists generated from academic abstracts, GPT-4 achieved 66.00% title accuracy, compared with 42.00% for GPT-3.5 and 36.98% for GPT-3; its relevance advantage was not clear. [c1] [c2] [c3] [c4] [c5]
Why it matters
For businesses using AI to draft research-backed content, a real source title is only one checkpoint—not proof that a reference is complete or useful.
| Model | Title accuracy |
|---|---|
| GPT-3 | 36.98% of evaluated citation titles; abstract-to-list task, CHI + EMNLP [c2] [c3] |
| GPT-3.5 | 42.00% of evaluated citation titles; abstract-to-list task, CHI + EMNLP [c2] [c3] |
| GPT-4 | 66.00% of evaluated citation titles; abstract-to-list task, CHI + EMNLP [c2] [c3] |
What to try
BLURSOR’s practical interpretation
Our interpretation: keep source-existence checks separate from editorial checks of whether a source supports the intended claim.
Study boundary
This was an exploratory academic-writing evaluation, not a test of business discovery or current retrieval-enabled AI search.
Do not replace review with a model upgrade
If your business uses AI to assemble references for an article, report or white paper, do not treat a model upgrade as a substitute for checking the evidence. This study supports a narrower conclusion: newer models can recommend more real paper titles without clearly recommending more relevant sources.
That distinction matters when deciding whether AI-assisted content is ready to publish. A reference can pass an existence check yet still be a poor fit for the document. The researchers examined academic citation recommendations, however—not marketing content, customer trust or brand visibility.
What the researchers actually tested
Courtni Byun, Piper Vasicek and Kevin Seppi compared GPT-3 text-davinci-003, GPT-3.5-turbo and GPT-4 using 20 target papers: ten from CHI, a human–computer interaction venue, and ten from EMNLP, a natural language processing venue.
Three tasks asked models to list citations from an abstract, write a Related Works section discussing earlier research, or rewrite a discussion with supporting citations. After excluding citations without titles, the dataset contained 1,616 annotated citations.
“Title Accuracy” meant the percentage of recommended citations whose titles matched real papers. Author metrics and publication-year error were calculated only for references classified as real. These measures therefore do not establish complete-reference accuracy across everything the models generated.
Real titles improved; relevance remained uncertain
In the abstract-to-citation-list task, pooled title accuracy was 66.00% for GPT-4, 42.00% for GPT-3.5 and 36.98% for GPT-3. GPT-4’s results also differed by field: 54.00% for human–computer interaction and 78.00% for natural language processing. These are title-existence rates among evaluated citations, not percentages of fully correct or useful references.
The authors concluded that GPT-4 typically performed better on accuracy but did not clearly outperform earlier models on relevance. Their relevance measure was exploratory: it checked overlap between recommendations and the original papers’ bibliographies, rather than asking people whether each source was useful.
That proxy can miss worthwhile recommendations the original authors did not cite. The paper also reports no significance results establishing model differences in relevance, so the result should not be read as proof that the models were equally useful.
A realistic publishing interpretation
Our interpretation is to separate bibliographic verification from editorial judgment. The first asks whether the source exists and whether its details are correct. The second asks whether its contents support the claim you intend to publish. The study measured title existence and bibliography overlap; it did not test this proposed business workflow.
- Check the original source rather than accepting a formatted reference as verification.
- Read the relevant passage before using a citation to support a business claim.
- Assess your actual subject area and task; the reported results varied across fields.
What this cannot tell your business
The study does not show how to earn citations in AI answers, whether an AI system will recommend your business, or how current retrieval-enabled products perform. Its practical value is a warning about interpreting citation-quality metrics, not a discoverability tactic.
Verification depended on exact title matches on the first Google Scholar results page, potentially misclassifying real papers as fabricated. Some analysis groups had only 24 observations. Appendix B also says GPT-3 was excluded from two final prose-task results, although the tables display GPT-3 results for those tasks. That unresolved inconsistency limits claims of identical final-prompt comparisons.
The source and evidence
This Reference Does Not Exist: An Exploration of LLM Citation Accuracy and Relevance
Study limitations and disclosures
- The authors are affiliated with Brigham Young University. The paper reports no funding statement or commercial-interest disclosure; that absence does not establish that none existed.
- The evaluation covers 20 academic target papers and older GPT models, not commercial recommendations or current AI search. Citations without titles were excluded, and some long discussions were shortened.
- Relevance was measured through bibliography overlap, not human judgments of usefulness or direct checks of claim support. Relevant sources absent from the target bibliography could go uncounted.
- Some analysis groups contained only 24 observations. Reported significance tests do not establish model differences on relevance.
- First-page Google Scholar title matching could miss real papers. Page numbers, publication venues and URLs were not evaluated, and manual curation could leave undetected errors.
- Appendix B and the main tables conflict over GPT-3’s inclusion in two prose-writing tasks, limiting comparisons based on identical final prompts.
Evidence behind this briefing
[c1] The study sampled 20 academic papers, ten from each of CHI and EMNLP. Three tasks requested citations from an abstract, a generated Related Works section, or a rewritten discussion. Citations lacking titles were excluded, leaving 1,616 annotated citations.
Section 3.3 Dataset, page 2 (printed page 29); task definitions in Section 3.2; target-paper list in Appendix E · Read the study
[c2] Title accuracy measures the percentage of recommended citations with real paper titles. Author precision, author recall and publication-year error were calculated only for citations classified as real papers; they do not describe all generated citations. Year accuracy is an absolute error in years, not a percentage.
Section 4.1 Accuracy, page 3 (printed page 30) · Read the study
[c3] For the Abstract→Citations List task, pooled title accuracy was 66.00% for GPT-4, versus 42.00% for GPT-3.5 and 36.98% for GPT-3. GPT-4 scored 54.00% for HCI and 78.00% for NLP. These are citation-title existence rates, not complete-reference accuracy or relevance; the table provides no confidence intervals for these percentages.
Table 1, Title Accuracy rows; first task block, HCI/NLP/Total columns, page 3 (printed page 30) · Read the study
[c4] Higher citation accuracy did not clearly translate into higher relevance: GPT-4 typically surpassed earlier models on accuracy but did not clearly outperform them on relevance. This is an observed model comparison, not evidence that upgrading a model causally improves citation usefulness.
Section 5 Conclusion, page 4 (printed page 31); qualified comparisons in Sections 4.1–4.2 · Read the study
[c5] Relevance was an exploratory bibliography-overlap proxy, not a human judgment of usefulness or a user-preference measure. Potentially relevant recommendations absent from the original authors’ bibliographies could therefore go uncounted.
Section 3.3 Dataset, page 3 (printed page 30); relevance definitions in Section 4.2 · Read the study
[c6] Despite 1,616 total citations, several analysis groups contained fewer than 30 observations, with a minimum of 24. Reported significance tests were limited to title accuracy between HCI and NLP; the paper does not supply significance results establishing model differences on relevance.
Section 6 Limitations, page 5 (printed page 32); reported tests in Table 3 · Read the study
[c7] Citation verification relied on exact title matches on the first page of Google Scholar results. Consequently, a citation classified as fabricated might instead be a real paper missed by that search procedure.
Section 6 Limitations, page 5 (printed page 32); classification procedure in Section 3.3, page 3 · Read the study
[c8] Appendix B discloses that final prompts for two prose-writing tasks exceeded GPT-3’s context length and says GPT-3 was not ultimately included in those results, although Tables 1–3 display GPT-3 results for those tasks. This unresolved reporting inconsistency limits claims that all three models received identical final prompts.
Appendix B Prompt Engineering, page 7 (printed page 34), contrasted with Tables 1–3 · Read the study