The practical brief
The finding
Across three search-enabled models, 83.6% of evaluated bibliographic fields were correct, but only 50.9% of reference entries had every evaluated field correct. [c1] [c2] [c4] [c6]
Why it matters
For businesses publishing research-backed content, a plausible reference is not the same as an accurately recorded source.
| Measure | Search-enabled baseline | Two-stage clibib workflow |
|---|---|---|
| Field accuracy | 83.6% of 23,103 evaluable bibliographic fields across three models | 91.5% of evaluable bibliographic fields across three models on the 931-paper benchmark [c1] [c2] [c4] |
| Fully correct entries | 50.9% of 2,793 reference entries across three models | 78.3% of reference entries across three models on the 931-paper benchmark [c1] [c2] [c4] |
What to try
BLURSOR’s practical interpretation
Our interpretation: separate finding a source from verifying its reference details, and retain editorial review before publication.
Study boundary
This tests scientific reference metadata, not brand discoverability or claim support. The correction workflow received additional identifying inputs, and lost underlying records prevent exact numerical reproduction.
Decide who checks references before publication
If your business uses AI to prepare research-backed articles, reports or white papers, do not treat web search as automatic reference verification. This study identifies a specific publishing risk: models can produce convincing-looking references with incorrect or missing details.
The researchers tested BibTeX, a structured format for recording titles, authors, dates and other citation metadata. Their results concern those details—not whether a cited paper supports your marketing claim.
The test supplied a known paper to find
For each of 931 scientific papers, researchers supplied its title and first author, deliberately omitting identifiers such as DOI—a persistent document identifier. GPT-5, Claude Sonnet 4.6 and Gemini 3 Flash each generated one reference per paper with vendor search enabled: 2,793 entries across four domains.
Across 23,103 evaluable fields, accuracy was 83.6%. Only 50.9% of entries were fully correct under the scoring rules. Several accurate details could therefore conceal an error elsewhere.
Field accuracy was 92.7% for popular papers, 82.0% for low-citation papers and 65.0% for recent papers. This association does not establish that gaining citations improves accuracy. Opaque search tools prevented researchers from distinguishing retrieved details from model memory.
Correction improved results, but inputs also changed
The authors tested clibib, their open-source tool for retrieving bibliographic records without generating metadata fields. In the two-stage workflow, models first searched, then revised references using retrieved records.
Crucially, the client-side lookup used paper metadata in priority order: URL, DOI, then title. The baseline prompt supplied only title and first author and deliberately omitted DOI. The comparison therefore does not isolate workflow architecture under identical identifying inputs.
Across all three models, reported field accuracy reached 91.5%, and fully correct entries reached 78.3%. The regression rate—previously correct fields becoming incorrect—was 0.8%. These are encouraging correction results, not proof that separation alone caused the gains.
Coverage remained incomplete. In single-stage tests involving only GPT-5 and Claude, 11.1% of 2,617 lookup calls returned no record. A returned record was not itself proof of accurate metadata.
Treat verification as an editorial safeguard
Our interpretation is to make reference verification a distinct editorial responsibility. This study offers no evidence that changing a business website earns more AI citations, or that these errors affect customer trust or sales.
For newer research, leave unresolved details unresolved rather than asking a model to fill gaps. The observed recency disadvantage warrants caution without making every recent source unreliable.
- Check the intended paper and its details against a publisher or authoritative bibliographic record.
- Review the full author list: primary scoring checked only the first author’s surname.
- Separately check whether the source supports the claim you plan to publish.
Reported gains have important verification limits
This was a single-point-in-time evaluation of three models, using English-language venues and vendor-recommended generation settings. Open-ended reference generation, other disciplines, languages, later models and different settings may behave differently.
An unrecoverable hardware failure destroyed the original benchmark records, raw outputs, complete annotations and analysis tables. Surviving materials allow method inspection, but not exact numerical reproduction. Ground truth also shared Zotero infrastructure with clibib, limiting evaluation independence.
Both authors list the University of Pennsylvania and evaluated their own tool. Funding came from DARPA SciFy and ODNI/IARPA BENGAL; the conclusions do not necessarily represent government policy.
The source and evidence
BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation
Study limitations and disclosures
- Original benchmark records, raw model outputs, complete annotations and analysis tables were lost in a hardware failure. Surviving materials permit method inspection, not exact numerical reproduction.
- The baseline prompt supplied title and first author, deliberately omitting DOI and other identifying metadata. Two-stage client-side lookup prioritized URL, DOI, then title. The mitigation comparison therefore does not isolate workflow architecture under identical identifying inputs.
- The evaluation covered one known-item retrieval strategy, four scientific domains and three models at one point in time. It used English-language venues and vendor-recommended generation temperatures, without a temperature-zero comparison. Other tasks, publication norms, languages, later models and settings may yield different results. Purpose-built alternatives such as CheckIfExist were not compared.
- Ground truth and clibib share Zotero infrastructure. Independent database checks covered only some papers; materials science relied on clibib alone. This limits the independence of the tool evaluation.
- Primary author scoring checked only the first author’s surname. Human–AI judging agreement was 80.4% on 521 checked fields, with systematic DOI and entry-type disagreements. Search provenance was unavailable, preventing attribution of details to retrieved pages versus model memory.
- Single-stage tool tests covered only GPT-5 and Claude, whereas two-stage results covered all three models. Lookup coverage was incomplete, and returning a bibliographic record did not establish its accuracy.
- Both authors, Delip Rao and Chris Callison-Burch, list the University of Pennsylvania as their affiliation. They evaluated their own open-source, MIT-licensed clibib tool; this is a relevant developer interest, not evidence of an additional commercial relationship. Funding came from DARPA SciFy and ODNI/IARPA BENGAL. The conclusions are the authors’ and do not necessarily represent official government policy.
- The authors used Gemini-Pro 3.0 and Claude Sonnet 4.5 for proofreading and plotting, with all outputs manually reviewed. They also used clibib to generate the paper’s BibTeX entries, which they manually reviewed for accuracy.
Evidence behind this briefing
[c1] The benchmark uses single-turn known-item retrieval: title and first author are supplied, other identifying metadata omitted. Each of 931 papers is queried once per model, producing 2,793 entries across GPT-5, Claude Sonnet 4.6, and Gemini 3 Flash with vendor search enabled. It covers four scientific domains and popular, low-citation, and recent tiers.
Section 3.1; model and sample scope in Sections 3.2–3.3, Tables 1–2 · Read the study
[c2] Baseline field accuracy is 83.6% across 23,103 evaluable fields, while 50.9% of entries have every evaluated field correct. These are bibliographic matching metrics, not citation preference or general factual accuracy; primary author matching checks only the first author's surname.
Section 4 opening; author matching qualifier in Sections 3.6 and 4.1 · Read the study
[c3] Popular papers score 92.7%, low-citation papers 82.0%, and recent papers 65.0% on field accuracy: a 27.7-percentage-point popular-to-recent gap. This is an observational tier comparison, not proof of a causal memorization mechanism. Some 2025 AI papers fall within Claude's reported knowledge window, and Appendix E.3 attributes the common 2025–2026 decline to indexing recency.
Section 4.4, Figure 1; cutoff qualification in Appendix E.3 and Section 3.2 · Read the study
[c4] Two-stage clibib integration across all three models raises reported field accuracy to 91.5% and fully correct entries to 78.3%, correcting 52.3% of baseline errors with 0.8% regression. The paper reports the gain as +8.0 percentage points despite rounded endpoints of 83.6% and 91.5%. Single-stage covers only GPT-5 and Claude; on those same models, two-stage gains over single-stage are significant by McNemar's test (p<10^{-4}).
Section 5.4, Table 4; integration protocol in Section 5.2 · Read the study
[c5] Ground truth shares Zotero infrastructure with the evaluated mitigation tool. A restricted analysis covers 694 papers in domains targeted for independent cross-validation, but that does not mean every paper was independently validated: Appendix E.6 reports DBLP matches for only 22/250 quantum-computing papers and 216/247 AI papers, and PubMed matches for 148/250 medicine papers. Materials science relies on clibib alone.
Appendix E.6; circularity analysis in Section 6.2 · Read the study
[c6] The original benchmark records, raw model outputs, complete field annotations, and analysis tables were lost in an unrecoverable hardware failure. Surviving code, prompts, rubrics, and aggregate figures permit method inspection but not exact numerical reproduction; reported results must be briefed with this material disclosure.
Artifact Availability · Read the study
[c7] Across GPT-5, Claude 4.6 and Gemini 3 Flash, pooled field-level accuracy declines from 92.7% for popular papers to 82.0% for low-citation papers and 65.0% for recent post-cutoff papers. Fully correct entry rates also decline. This is a citation-tier association, not evidence that increasing citations causally improves AI accuracy. Table 12 defines field-level accuracy over evaluable fields and supplies 95% bootstrap confidence intervals from 2,000 cluster-resampled iterations.
Appendix F.9, Table 12 · Read the study
[c8] Reported baseline author accuracy uses only the first author's last name, matching a prompt that provides title and first author. Requiring all authors' last names lowers pooled accuracy from 90.7% to 79.6%, an 11.1-percentage-point drop; ordered matching yields 79.5%. This lenient baseline inflates apparent author completeness.
Appendix F.8, Table 11 discussion · Read the study
[c9] The LLM judge and human annotator were compared on 521 Stage 2 fields. Overall agreement was 80.4% with kappa 0.671; systematic entry-type and DOI disagreements matter. Excluding both fields increases kappa to 0.838, so that higher agreement should not be presented as applying to all fields.
Appendix F.5, Table 10 and preceding discussion · Read the study
[c10] Vendor search tools are opaque: the study cannot attribute output fields to transcription from retrieved pages versus reconstruction from model memory. Search-source counts also reflect vendor implementations rather than a researcher configuration choice.
Appendix F.10, Search depth details · Read the study
[c11] For GPT-5 and Claude only, single-stage clibib augmentation corrected 1,277 of 2,944 baseline field errors (43.4%). Missing fields were corrected more often (924/1,877; 49.2%) than fabricated (57/315; 18.1%) or substituted fields (28/188; 14.9%). Gemini was excluded because its API could not combine GoogleSearch grounding with function declarations in one request.
Appendix G.4, Table 13; model exclusion disclosed in G.1 · Read the study
[c12] Single-stage clibib lookup coverage is not universal: 88.9% of 2,617 calls returned valid BibTeX, while 11.1% returned not_found. Success was lowest in AI (81.8%); recent papers had a 14.7% not-found rate. These are lookup-call outcomes, not verified metadata-accuracy rates. Unavailable records and wrong-paper substitutions remain residual failure modes.
Appendix G.7, Table 14 discussion; single-stage GPT-5 and Claude only · Read the study
[c13] Both authors, Delip Rao and Chris Callison-Burch, list the University of Pennsylvania as their affiliation. This establishes their stated institutional affiliation, not additional commercial interests.
Title-page author blocks · Read the study
[c14] The authors disclose using Gemini-Pro 3.0 and Claude Sonnet 4.5 for proofreading and plotting, with all outputs manually reviewed. They also used clibib to create all BibTeX entries and manually reviewed them for accuracy.
Generative AI Use Disclosure · Read the study
[c15] The evaluation used one known-item retrieval strategy, four scientific domains, and three models at one point in time. Open-ended generation and fields with different publication norms may behave differently; the results do not establish performance for later model versions.
Limitations · Read the study
[c16] The study used vendor-recommended temperatures rather than a temperature-zero comparison, covered English-language venues only, and did not compare against purpose-built tools such as CheckIfExist. Results under other generation settings or in other languages are not established.
Limitations · Read the study
[c17] The research discloses DARPA SciFy funding and partial support through ODNI/IARPA's BENGAL program, with a disclaimer that the authors' views do not necessarily represent official government policies.
Acknowledgments · Read the study
[c18] An unrecoverable hardware failure destroyed the original 931-paper benchmark records, raw model outputs, complete field-level annotations, and underlying analysis tables before archival. Surviving materials permit method inspection but not exact reproduction of the reported numerical results.
Artifact Availability · Read the study