The practical brief

The finding

In a 40-question medical-answer audit, direct PubMed identifier instructions produced perfect citation mapping but high automated source-support failure rates; GPT-5 nevertheless produced more supported claims per answer than GPT-4o. [c2] [c4] [c6] [c13]

Why it matters

For businesses publishing health-related content, confirming that a reference exists does not establish that it supports the accompanying claim.

Resolvable references and supported output are different measures
Audit measureGPT-4oGPT-5
Direct PMID citation mapping100% of direct-PMID citations mapped in the 40-question audit100% of direct-PMID citations mapped in the 40-question audit [c2] [c3]
Source-support failure with direct PMID instructions96.3% of successfully mapped claim–citation pairs in the medical-answer audit85.7% of successfully mapped claim–citation pairs in the medical-answer audit [c4] [c9]
Source-support failure across five conventional citation styles42.8%–55.8% of successfully mapped claim–citation pairs in the medical-answer audit, varying by style44.9%–53.0% of successfully mapped claim–citation pairs in the medical-answer audit, varying by style [c4] [c9]
Average source-supported claims per answer0.2–3.2 supported claims per answer across seven citation instructions in the 40-question audit2.6–8.3 supported claims per answer across seven citation instructions in the 40-question audit [c13] [c15] [c16] [c17]
Results from the 40-question medical-answer audit. Support was assessed automatically using PubMed titles and abstracts; partial agreement counted as support. Human baseline: 10.3 supported claims per answer. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: separate reference-resolution checks from claim-support review, using automated verification to prioritize review rather than approve publication.

Study boundary

This biomedical study did not test commercial discoverability or marketing outcomes; support assessments used article titles and abstracts.

The publishing risk extends beyond broken references

If your business uses AI to draft health-related articles, the decision is whether checking citations is enough before publication. This study suggests it is not: a reference can resolve to a real biomedical article without supporting the attached claim.

Researchers developed Med-V1, small language models that assess evidence attribution—whether a source supports an assertion. Across five biomedical datasets containing 14,274 claim–article pairs, two variants achieved macro-average accuracies of approximately 72.8% and 73.2%, versus 51.1% and 51.5% for their base models. Macro-average accuracy weights each dataset equally.

Performance was comparable to the tested frontier models. That made Med-V1 useful for auditing citations, not an infallible substitute for editorial judgment.

Evidence [c1] [c6]

Citation validity, support rate and supported yield differed

The audit used the first 40 medical questions from MedAESQA. GPT-4o and GPT-5 answered under seven citation instructions. GPT-5 extracted claim–citation pairs, a PubMed matcher resolved references, and Med-V1 checked claims against mapped articles’ titles and abstracts.

Direct PMID instructions—requesting PubMed article identifiers—produced perfect mapping. Yet among successfully mapped pairs, the verifier classified 96.3% for GPT-4o and 85.7% for GPT-5 as content hallucinations. Across five conventional citation styles, rates were 42.8%–55.8% and 44.9%–53.0%, respectively.

Here, “content hallucination” means no partial or full support was found in the cited source; it does not independently establish factual falsity. The study likewise defined citation fabrication as PubMed mapping failure, not proof that a reference exists nowhere.

Similar conditional failure rates did not mean equal supported output. GPT-5 produced an average of 2.6–8.3 source-supported claims per answer across citation instructions, versus GPT-4o’s 0.2–3.2; both were below the human baseline of 10.3. GPT-5’s longer answers therefore delivered more supported claims, despite similar failure rates under conventional styles. These counts remain automated title-and-abstract assessments, not full-text or medical-truth judgments.

Evidence [c2] [c3] [c4] [c9] [c13] [c17] [c20]

A realistic interpretation: review the claim–source relationship

For biomedical publishers, our interpretation is to treat reference validity and evidentiary support as separate review questions. A working identifier answers only the first.

A second audit illustrates the distinction beyond AI-generated answers. Researchers manually reviewed 100 selected, model-flagged contradictions in clinical guidelines and validated 32 misattributions by consensus. This selected sample does not estimate overall guideline error rates or demonstrate actual public-health harm.

  • Confirm that a reference resolves to the intended article, then inspect whether it supports the claim’s exact wording.
  • Use automated flags as review priorities: abstracts can omit supporting evidence available in full text.
  • Do not assume changing citation style will improve public AI visibility; that outcome was not tested.

Evidence [c5] [c6] [c8] [c11]

What the study cannot tell your business

The study did not evaluate end-to-end retrieval, commercial content, brand recommendations, conversions or customer trust. It also did not assess evidence quality, study design, risk of bias or generalizability. Source support is not the same as sound evidence.

Zero-shot evaluation meant no benchmark-specific fine-tuning, not complete separation of training and evaluation articles. Synthetic-label validation covered only 100 consensus-filtered instances, and training-design contributions were not isolated. The small citation audit used automated extraction and verification; its content-hallucination denominator excluded unmapped references.

The authors disclose NIH intramural and grant support, government-work status for NIH authors’ contributions, and no competing interests. Their conclusions do not necessarily reflect NIH or government views. Models, data and code are publicly available, but business deployment costs and suitability remain unestablished.

Evidence [c6] [c7] [c9] [c12]

The source and evidence

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, and Zhiyong Lu · 2026-03-05

Study limitations and disclosures

  • The source is restricted to biomedical evidence attribution; it does not measure brand visibility, citation selection for commercial content, or marketing outcomes.
  • Citation fabrication is operationalized as failure to map to PubMed, rather than independent confirmation that a cited work does not exist.
  • Figure 4 error bars are 95% confidence intervals estimated using 2,000 bootstrap iterations; numeric interval endpoints are not supplied in this chunk.
  • Benchmark convergence is not presented as general task saturation: ambiguous claims, label noise and binary-question reformulation artifacts may limit discrimination.
  • The error analysis sampled 20 misclassified predictions per dataset per variant; its category proportions describe the reviewed sample, not the benchmark-wide error distribution.
  • Article overlap existed: 558 of 6,868 resolved unique benchmark PMIDs appeared in training. Excluding overlapping articles changed macro-average accuracy from 0.7276 to 0.7266 for Med-V1-L3B and from 0.7316 to 0.7261 for Med-V1-Q3B.
  • A validated guideline misattribution means the cited source does not support the statement, not necessarily that the statement is universally incorrect.
  • Model checkpoints, synthetic training data, evaluation resources and training code are disclosed as publicly available. Synthetic-data development took approximately four elapsed months, not four months of active compute; SFT on eight H200 GPUs required on the order of tens of GPU-days.
  • The benchmark's primary metric is macro-average accuracy across five datasets, weighting datasets equally rather than weighting every instance equally.
  • MedAESQA preprocessing discards missing or invalid PMIDs and merges neutral and not relevant into NEI; Boolean QA datasets are repurposed into verification tasks rather than originally annotated for that task.
  • Error analysis samples 20 ground-truth disagreements per component dataset per model variant: 100 cases each for Med-V1-L3B and Med-V1-Q3B, 200 total. Consensus categories are assigned hierarchically, so these counts do not represent all benchmark cases.
  • GPT-5 performs claim-citation extraction. Human expert answer statistics use the same extraction pipeline, and average mapped PMID value is explicitly a proxy for recency.
  • PubMed mapping failure is the study's fabrication criterion, not independent proof that a reference does not exist; the content-hallucination denominator excludes unmapped citations.
  • Data, code, and both Med-V1 models are disclosed as publicly available. Listed API models are accessed through Microsoft Azure, with temperature 0 wherever applicable.
  • The supported-claim yield is an automated title-and-abstract attribution measure, not a direct measure of medical truth, evidence quality, or full-text support.
  • The reported conditional hallucination rates exclude citations that failed PMID mapping; supported claims per answer and conditional support rates must not be treated as interchangeable.
  • The comparison covers the first 40 MedAESQA questions and seven citation instructions. Numerical confidence-interval bounds are not supplied in this chunk.
  • The support assessment uses an automated verifier with titles and abstracts, rather than full-text verification.

Evidence behind this briefing

[c1] On five biomedical datasets totaling 14,274 claim–article pairs, the synthetic-data-trained 3B Med-V1 variants achieved zero-shot macro-average accuracies of approximately 0.728 and 0.732, versus 0.511 and 0.515 for their respective base models. The reported 42.5% and 42.0% gains are relative accuracy improvements, not percentage-point gains; frontier-model averages ranged from 0.717 to 0.736.

2.3 Med-V1 Closes the Performance Gap between Frontier and Lightweight LLMs · Read the study

[c2] The medical-answer citation audit used only the first 40 MedAESQA questions and seven citation instructions. GPT-5 extracted claim–citation pairs from GPT-4o and GPT-5 answers; the PubMed single citation matcher resolved citations, and Med-V1-L3B assessed claims using mapped articles' titles and abstracts. Mapping failure operationalized citation fabrication; content hallucination was assessed only among successfully mapped pairs.

2.5 Use Case 1: Detecting LLM Hallucinations with Med-V1 · Read the study

[c3] In this 40-question medical-answer audit, NLM, AMA and Vancouver citations had higher successful PMID-mapping rates (61.4%–83.3%) than APA and MLA citations (44.6%–50.4%). Direct PMID citations mapped perfectly, but mapping alone did not establish evidentiary support. Human answers averaged 10.3 claims; GPT-4o generated 5.1–7.4 and GPT-5 18.6–36.3 per answer.

2.5 Use Case 1; Figure 4b–c · Read the study

[c4] Among successfully mapped citations, verifier-defined content hallucination rates for five standard citation styles were 42.8%–55.8% for GPT-4o and 44.9%–53.0% for GPT-5. Direct PMID instructions yielded 96.3% and 85.7%, respectively. Partial support counted as supported; unsupported, neutral and contradictory scores counted as hallucinations. These are automated source-support assessments, not independently established factual-falsity rates.

2.5 Use Case 1; Figure 4e · Read the study

[c5] The guideline audit covered 6,152 publicly available PubMed Central guideline articles published in 2015–2025 and approximately 57,000 single-citation statement–source pairs. Of 100 flagged contradictions manually reviewed with equal sampling of partial and strong contradictions, 32 were consensus-validated misattributions; 17 of these 32 concerned treatment effectiveness. This selected validation sample does not estimate overall guideline misattribution prevalence or demonstrate actual public-health harm.

2.6 Use Case 2: Identifying High-Stakes Misattributions with Med-V1; Figure 5b–c · Read the study

[c6] The study primarily evaluates standalone claims against PubMed titles and abstracts, not full-text evidence or end-to-end retrieval. It does not assess evidence quality, study design, generalizability or risk of bias. Synthetic-label validation covered only 100 consensus-filtered instances; aggregation optimality, separate confidence scores and independent effects of individual-model versus explanation-bearing supervision were not established.

3 Discussion, limitations paragraph · Read the study

[c7] Zero-shot evaluation does not guarantee article-level separation from training. Overlap is checked by resolved PMIDs in either training role; unresolved or ambiguous PMIDs are not treated as non-overlapping.

Train-test article-overlap analysis · Read the study

[c8] The citation-instruction case study uses the first 40 MedAESQA medical questions, GPT-4o and GPT-5, and seven citation formats. This is a specific medical-answer test, not a general search-visibility experiment.

4.6 Detecting LLM Hallucinations with Med-V1, opening paragraph · Read the study

[c9] Citation fabrication is operationalized as failure to map to PubMed, with the top-ranked citation-matcher result retained. Content hallucination is assessed by Med-V1-L3B against titles and abstracts only, conditional on successful PMID mapping; partial agreement counts as support.

4.6 Detecting LLM Hallucinations with Med-V1, citation mapping and scoring · Read the study

[c10] The guideline study starts with 6,152 free full-text practice-guideline PMIDs from 2015–2025, filters to about 57k claim-citation pairs, and samples 100 model-flagged disagreements equally across partial and strong contradiction. This selected sample is not a random sample of all guideline citations.

4.7 Identifying High-Stakes Misattributions with Med-V1, corpus and sampling · Read the study

[c11] Guideline validation acknowledges that full-text support may be absent from abstracts. Two annotators independently review the 100 selected cases and adjudicate by consensus, distinguishing insufficient abstract information from model error and validated misattribution.

4.7 Identifying High-Stakes Misattributions with Med-V1, annotation procedure · Read the study

[c12] The authors disclose NIH intramural and grant support, a government-work designation for NIH authors' contributions, a disclaimer of institutional endorsement, and no declared competing interests.

Acknowledgements; Competing interests · Read the study

[c13] GPT-5 generated 2.6–8.3 source-supported claims per answer versus GPT-4o's 0.2–3.2, while both remained below the human baseline of 10.3. Thus, comparable conditional hallucination rates do not mean comparable absolute supported-claim yield.

Section 2.5, final results paragraph; Figure 4g · Read the study

[c14] For five standard citation formats, conditional content hallucination rates among successfully PMID-mapped citations were similar: 42.8%–55.8% for GPT-4o and 44.9%–53.0% for GPT-5. Support includes partial or strong agreement; these rates are distinct from supported claims per answer.

Section 2.5, content hallucination results paragraph; Figure 4e · Read the study

[c15] The comparison used the first 40 MedAESQA medical questions under seven citation instructions.

Section 2.5, second paragraph · Read the study

[c16] Support was assessed automatically by Med-V1-L3B using mapped PubMed titles and abstracts, after GPT-5 extracted claim-citation pairs; it was not a full-text or human adjudication of these answers.

Section 2.5, third paragraph · Read the study

[c17] Figure 4 presents average supported-claim counts and error bars representing 95% confidence intervals estimated from 2,000 bootstrap iterations; the supplied prose does not give numerical interval bounds.

Figure 4 caption, panel g and statistical notes · Read the study

[c18] The comparison tested GPT-4o and GPT-5 on the first 40 medical questions from MedAESQA under seven citation-instruction formats.

4.6 Detecting LLM Hallucinations with Med-V1, opening paragraph · Read the study

[c19] Source support was assessed automatically by Med-V1-L3B against mapped PubMed titles and abstracts. Partial or strong agreement counted as support; the content-hallucination rate was conditional on successful PMID mapping, not a measure of absolute supported-claim yield per answer.

4.6 Detecting LLM Hallucinations with Med-V1, citation mapping and claim verification paragraphs · Read the study

[c20] The study reports supported-claim proportions and supported-claim counts separately, alongside the conditional content-hallucination rate. Human expert statistics use MedAESQA expert answers processed through the same claim-extraction pipeline.

4.6 Detecting LLM Hallucinations with Med-V1, final paragraph · Read the study