The practical brief

The finding

In a clinical-guideline test, claude-opus-5 attached verbatim-compliant quotations to 98.0% of its 658 claim sentences, but the quotations fully supported only 37.1%. Compliance included normalized and elided matches, not just exact reproduction. [c1] [c2] [c11] [c13] [c15]

Why it matters

A quotation matching a source is not sufficient evidence that an AI-generated statement accurately represents it.

Source matching versus complete support
MeasureResult and denominator
At least one verbatim-compliant quotation98.0% of 658 clinical claim sentences [c13] [c14] [c15]
Fully supported by accompanying quotations37.1% of 658 clinical claim sentences [c2] [c15]
Fully supported or supporting evidence recovered85.1% of 658 clinical claim sentences after the recovery probe [c3] [c15]
All rows concern claude-opus-5 in the controlled clinical-guideline test. Matching includes exact, normalized, and elided quotations; recovery was a separate post-generation search. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: before publishing sourced AI content, check whether the evidence supports every material detail, not merely whether the quotation exists.

Study boundary

This controlled clinical verifiability test did not measure business discoverability, customer trust, or overall answer accuracy.

A real quotation should not automatically clear publication

If your business uses AI to draft sourced explanations, the publication decision should go beyond checking that a quotation exists. A real passage can support part of a statement while leaving an important detail unsubstantiated.

This study investigated that gap in clinical question answering. Its business relevance is a verification principle, not proof of a marketing outcome. Being quoted in an AI answer and having your source accurately represented are different questions; this experiment addresses quotation support, not discoverability.

Evidence [c2] [c10] [c11]

The test separated source matching from claim support

Researchers compared twelve language models from five vendors on 222 synthetic clinical questions designed to be answerable from four guidelines. Each generation model received the same already-selected guideline sections and could not retrieve additional material.

The main analysis linked each citation to the nearest preceding claim sentence. “Verbatim-compliant” meant a quotation passed exact, normalized, or elided matching. Normalization allowed formatting changes, including removal of reference markers, bullets, and table markup. Elided matching allowed omitted text between at least two source-ordered fragments of ten or more characters.

A separate AI judge assessed whether the accompanying quotations supported the claim, using the question as context rather than evidence. Two blinded human annotators checked a stratified sample of 70 cases. Strong agreement supports rating consistency, not clinical correctness. A closing footnote also states, “Both are authors of this paper,” but the supplied text does not identify whom “Both” refers to.

Evidence [c1] [c4] [c9] [c10] [c12] [c13] [c14]

Matching quotations often left details unsupported

Under that inclusive matching definition, claude-opus-5 supplied at least one compliant quotation for 98.0% of its 658 claim sentences. Yet only 37.1% of those same claims were fully supported by their accompanying quotations. The gap largely reflected partial support rather than a complete absence of evidence.

This does not mean the remaining claims were medically false. An automated recovery probe searched the guideline sections the model had already read for missing supporting passages. Counting fully supported claims plus recovered evidence raised the corresponding share to 85.1%.

A realistic interpretation is that evidence selection can fail even when useful evidence is available. Recovery suggests a possible checking step; it does not demonstrate faster reader verification or better performance in a deployed interface.

Evidence [c2] [c3] [c7] [c15]

Use the finding as an editorial distinction

Our practical interpretation is to separate three questions: does the quotation match the source, does it support the whole statement, and is the statement appropriate for publication? The paper tests the first two much more directly than the third.

For an important claim, a reviewer could compare its material details with the evidence, then add missing support or narrow the wording. This is an editorial application, not an experimentally validated tactic for business content.

  • Treat partial support as a reason to inspect the claim, not automatically label it false.
  • Do not infer better AI visibility, recommendations, customer trust, or sales from these results.
  • Keep proposed review workflows conditional: business adoption, review costs, and acceptable delays were not measured.

Evidence [c2] [c3] [c10] [c11]

The source and evidence

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst · 2026-09-14

Study limitations and disclosures

  • Synthetic questions, four clinical guidelines, frozen retrieval, and one system prompt limit generalization to real users or marketing tasks.
  • AI judges assessed quotation support; human validation covered 70 cases. The study evaluates verifiability, not whole-system performance or clinical correctness.
  • Recovery input was capped at 380,000 characters. Of 2,466 probes, 103 reached that cap, potentially excluding available evidence.
  • The paper discloses, “Both are authors of this paper.” The supplied text leaves the two people unidentified, so this disclosure cannot reliably be attributed to the human annotators or other participants.
  • Verbatim compliance includes normalized and elided source matches, not only character-for-character quotations; it should not be read as an exact-match rate.

Evidence behind this briefing

[c1] The comparison covers twelve LLMs from five vendors answering 222 synthetic clinical queries designed to be answerable from four clinical guidelines. Retrieval output was frozen across generation models; the main analysis attributes each citation to the nearest preceding claim sentence (k=1).

Section 1, primary contributions; Sections 3.1–3.2, task and attribution-window details · Read the study

[c2] In this harness, claude-opus-5 supplied at least one verbatim-compliant quote for 98.0% of its 658 claim sentences, but only 37.1% of all those claims were fully supported by their accompanying quotes. This is a quotation-support result, not a finding that the remaining claims were medically false. Among its 645 claims with compliant quotes, strict support was 37.8% and lenient support was 99.7%.

Section 4, results overview; Table 1, claude-opus-5 row; Table 3, claude-opus-5 row · Read the study

[c3] An automated recovery probe searched the sections already read for missing evidence supporting partially supported claims. For claude-opus-5, the share of all claims fully supported or recovered rose from 37.1% to 85.1%; several GPT models reached approximately 92%. These are post-generation evidence-recovery results, not improvements measured in user verification speed or a deployed answer interface.

Section 4, Finding 3: Most missing evidence is already in what models read · Read the study

[c4] Claim support was judged by deepseek-v4-flash and validated against two blinded human annotators on a stratified sample of 70 cases. Human–human Cohen’s kappa was 0.853; human–judge kappas were 0.858 and 0.810. Agreement with a second LLM judge over the full samples was 0.803. These agreement statistics validate rating consistency, not clinical correctness.

Section 4, Finding 2; Appendices A.12–A.13 · Read the study

[c5] The authors explicitly limit the evaluation to synthetic questions and one system prompt without prompt ablations.

Section 5, Limitations · Read the study

[c6] Recovery evidence had to pass the verbatim matcher, but the recovery judge’s read set was capped at 380,000 characters. Of 2,466 probes, 103 reached this cap and 13 of those returned no span; the cap could exclude available evidence from the recovery input.

Appendix A.14, Recovery from the Read Set · Read the study

[c7] Recovery used deepseek-v4-flash to search the guideline sections the generation model had read for support missing from its cited quotes, followed by verbatim source verification. The judge's read set was capped at 380,000 characters; 103 of 2,466 probes reached that cap, and 13 capped probes returned no span.

Appendix A.14, Recovery from the Read Set · Read the study

[c8] Among claims judged partially supported with an unsupported addition, verified recovery ranged from 87.0% to 97.7% across the 12 generation models. These percentages measure recovery of a verbatim-verified supporting span from the read set, not overall answer accuracy; model-specific denominators ranged from 54 to 383 cases.

Appendix A.14, Table 11 · Read the study

[c9] The frozen-context generation prompt restricted every generation model to already-retrieved clinical-guideline sections; models could not retrieve additional material. This is a controlled grounding test rather than an unrestricted AI-search setting.

Appendix B.3, Generation — System prompt: frozen-context evaluation · Read the study

[c10] The support judge was instructed to evaluate only the cited quote as evidence, treating the question as context rather than evidence. Consequently, quote-support verdicts do not establish whether a claim is supported elsewhere in the guidelines.

Appendix B.5, Claim Support Judge — Hard rules · Read the study

[c11] The authors explicitly limit the work to evaluating verifiability, rather than performance of the system as a whole.

Footnotes, footnote 1 · Read the study

[c12] The closing footnotes disclose majority voting over three sampled judgments, with extra rounds for ties; distinguish strict support rate over cited claims from CCR over all claims despite a shared numerator; and state that two unspecified people are paper authors. The supplied text does not identify the referents of that authorship disclosure.

Footnotes, closing three statements · Read the study

[c13] “Verbatim-compliant” does not mean only character-for-character reproduction. A quotation must appear in the cited section, but it can pass as an exact, normalized, or elided match. The citation-level Verbatim Compliance Rate is the fraction of citations whose quotes pass this check.

Section 3.2, Evaluation Framework and Metrics — Verbatim Compliance · Read the study

[c14] Normalized matching permits Unicode, case, quotation-mark, dash, and whitespace normalization, including removal of guideline reference markers, list bullets, and table markup. Elided matching accepts ellipsis-separated fragments occurring in source order, disregards fragments shorter than ten characters, and requires at least two remaining fragments.

Appendix A.6, Verbatim Match Tiers · Read the study

[c15] Under that inclusive matching definition, claude-opus-5 retained 98.0% of its 658 claim sentences at the verbatim stage, while 37.1% of those same 658 claims reached the fully supported certified-claim stage. These are claim-level funnel percentages, not citation-level exact-match rates.

Section 4, Table 1 — The claim funnel; claude-opus-5 row (Sent., Claims, Covered, Verbatim, Lenient, CCR (Strict), Recov., SCR) · Read the study