The practical brief

The finding

On 928 injected citation instances, image-aware verification achieved 93.0–93.7% accuracy against page-localization labels and outperformed compute-matched stronger OCR. [c1] [c2] [c13] [c14] [c20]

Why it matters

A clickable page reference is not sufficient evidence that a business document supports an AI-generated claim.

Verification success is not loss-free correction
Measure and scopeReported result
Verification accuracy93.0–93.7% across three model families on the 928-instance injected benchmark; page-localization labels are a support proxy [c2] [c14]
Citation precision after correction87.0–89.9% of retained citations across three model families in the injected mix, from a constructed 34.3% starting precision [c3]
Supported-claim retentionAmong 450 human-confirmed supported claims in the natural candidate-error pool: Gemini 94.7%; Claude 94.7%; GPT 88.0% [c20]
Uncached experimental costGemini injected benchmark, provider list prices: default verifier $0.024 per item; layout OCR $0.009 per item [c18] [c19]
Injected results use a constructed error mix and page-localization labels. Natural-pool retention is conditional on selected suspected errors. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: test citation checking on your own documents, including comparative accuracy, cost and supported content removed.

Study boundary

The study tests English business/legal document images, not public AI-search visibility or customer outcomes; page localization is a proxy for semantic support.

Check support before releasing cited answers

If your business uses AI to explain document packs to customers, decide how references will be checked before releasing answers. A real page number can accompany a claim that the page does not support.

This study examines verification inside a known document, not how businesses win citations in public AI search. Its relevance is to buying or building document-answering systems where the original evidence is available.

Evidence [c1]

What the verification test actually measured

AtomCite separates answers into claims, checks cited page images and applies a fixed correction policy. OCR—optical character recognition—extracts text from images; comparison systems relied on that text. Only the cited page can justify the original reference. Up to two neighboring pages in either direction can supply replacement proposals.

The main test used 928 validated, deliberately injected instances from English business/legal document images. Across Gemini, Claude and GPT, verification accuracy was 93.0–93.7% against deterministic page-localization labels: whether the cited page belonged to a reference page set.

That is a proxy for semantic support—whether the evidence actually establishes the claim—not direct proof. Separate human calibration found 7.2% disagreement between localization and semantic judgments, with reported interval [3.1%, 15.9%].

Evidence [c1] [c2] [c13] [c14] [c15]

Better verification did not mean loss-free correction

Against compute-matched OCR using the stronger MinerU engine, AtomCite improved accuracy by 4.9 percentage points for Gemini, 2.3 for Claude and 3.9 for GPT. The respective 95% confidence intervals were [3.0, 6.9], [0.3, 4.2] and [2.1, 5.8] percentage points. Counting excluded verdicts as errors preserved system ordering, but the Claude comparison lost statistical significance.

Correction raised citation precision—the share of retained citations supporting claims under benchmark labels—from a constructed 34.3% to 87.0–89.9%, retaining 91.8–92.5% of correct claims. The starting precision is not an estimate of ordinary business-document quality.

Repair changed references or removed claims. Only 69.3% of wrong-page instances had a reference page within the permitted window. In the selected natural candidate-error pool, correction retained 94.7% of 450 human-confirmed supported claims with Gemini and Claude, and 88.0% with GPT. Some supported content was therefore lost.

Human auditing also reversed the apparent ranking of flat OCR and image-aware verification: automatic labels had rewarded false alarms. Replacement pages for natural errors still lack complete item-level human scoring.

Evidence [c2] [c3] [c4] [c8] [c9] [c10] [c11] [c20]

Buy an accuracy–cost trade-off, not a guarantee

Our interpretation: evaluate citation support separately from answer fluency. The stronger-OCR accuracy comparison above is distinct from the layout-OCR cost ledger. In the Gemini injected test, the default verifier averaged $0.024 and 2.8 calls per item, versus $0.009 and 1.3 calls for layout OCR.

Those are uncached provider-list-price averages, not deployment estimates. Default-verifier costs came from two cache-disabled repeats; absolute prices differ by provider.

Separate transfer tests improved long-form response checking but worsened some single-fact and already-separated-claim evaluations. More checking steps are not automatically better. The study cannot establish customer trust, commercial return or discoverability gains.

  • Ask vendors about false alarms and supported content removed, not just overall accuracy.
  • Pilot on your document types and answer lengths with an explicit cost budget.
  • Review consequential removals and replacement pages. These safeguards are our interpretation, not proven business tactics.

Evidence [c1] [c2] [c8] [c11] [c12] [c17] [c18] [c19] [c20]

The source and evidence

AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents

Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos · 2026-09-05

Study limitations and disclosures

  • Testing covered English business/legal document images within known documents, not open-web retrieval, other scripts, born-digital PDFs, non-page citations or AI-search visibility.
  • Primary labels measure page localization, not semantic entailment. Human calibration found 7.2% disagreement in injected strata, with reported interval [3.1%, 15.9%]. Expanded reference-page sets also contained imperfect evidence matches.
  • Synthetic precision gains depend on a constructed mix and restricted repair window. Natural-pool correction retained 88.0–94.7% of 450 supported claims; replacement-page accuracy remains incompletely human-scored.
  • Natural results concern suspected-error candidates, not all citations. Some undisputed candidates retained automatic labels rather than receiving individual human validation; mixed annotator judgments were excluded.
  • Counting inconsistent or unparseable verdicts as errors preserved system ordering but reduced statistically significant headline comparisons from 18 to 16. The Claude comparison against compute-matched MinerU OCR and GPT comparison against a sequential chain lost significance. Transfer tests also showed failures on some short or already-separated inputs.
  • The experimental cost premium over layout OCR is not a customer deployment estimate. Gemini costs used uncached provider-list-price averages and two cache-disabled repeats for the default verifier; absolute prices differ across providers.

Evidence behind this briefing

[c1] AtomCite evaluates claims against supplied citations within the source document, not open-web retrieval. Verdict evidence is restricted to the cited page; a ±2-page window, capped at five pages, supplies correction proposals only.

Section 4, (2) Evidence; closed-world scope defined in Section 1 · Read the study

[c2] On DocCite-Syn's 928 validated injected instances, binary cited-page support accuracy was 93.7% for Gemini and 93.0% for Claude and GPT. Against compute-matched OCR using the stronger MinerU engine, paired gains were 4.9, 2.3, and 3.9 percentage points, with respective 95% document-cluster bootstrap CIs of [3.0,6.9], [0.3,4.2], and [2.1,5.8]. All 18 headline comparisons were Holm-significant, although counting excluded verdicts as errors reduced that to 16.

Section 1, results summary; Tables 1–2; Section 5, Statistics · Read the study

[c3] Conservative correction raised citation precision on the constructed synthetic mix from 34.3% to 87.0–89.9%, retaining 91.8–92.5% of correct claims. Reimplemented CiteFix reached 50.3% precision and adapted RARR 36.8%; CiteFix's 100% retention is not informative because it cannot remove claims. Repair reach is limited: 69.3% of wrong-page instances had a gold page within ±2 pages, and remaining cases fell to removal.

Section 5.7, Correction; Table 5 and its † footnote; Appendix A, DocCite-Syn templates · Read the study

[c4] Automatic page-localization labels reversed the apparent ranking of flat OCR and image-grounded verification in the natural candidate-error pool. The audit partitioned 2,468 candidates into 1,909 errors, 450 supported claims and 109 mixed-annotator exclusions. On the 2,359 retained instances, AtomCite detected 97.3–98.6% of errors with 5.4–12.3% false alarms; flat OCR detected 99.7–99.8% but falsely flagged 43.8–49.3% of supported claims. These rates are conditional on the candidate-error pool, not all citations; 1,554 undisputed candidates retained automatic labels rather than receiving individual human validation.

Section 5.6, Natural Errors; Table 4; Section 3, DocCite-Nat · Read the study

[c5] In a separate text-only transfer test, frozen-prompt AtomCite without its image channel or correction policy improved both Qwen2.5-7B and Llama-3-8B over same-backbone single-call judges on all five public benchmarks. Metrics differ across benchmarks: F1, balanced accuracy and accuracy. On RAGTruth, Qwen improved by 21.7 F1 points, 95% CI [18.4,25.0], to 71.2; this remained below reported task-fine-tuned detectors.

Section 5.8, Transfer to Public Hallucination Benchmarks; Table 6 · Read the study

[c6] Expanded evidence-page gold sets are imperfect: in a 50-QA human audit covering all 76 added pages, only 60.5% carried fully equivalent evidence, with a reported interval of [49.3,70.8]. Expansion trades precision for recall, so page localization is not equivalent to semantic support.

Appendix A, Gold-set expansion · Read the study

[c7] Layer-1 candidate-error prevalence across six generator configurations was 44.7–61.2%; these are automated candidate labels, not human-confirmed error rates. Denominators cover the current snapshot reproducing 2,392 of 2,468 pool rows, excluding 76 superseded rows without denominators. Neither frontier nor efficiency tier was uniformly cleaner.

Appendix A, Candidate-error prevalence per generator; Table A1 · Read the study

[c8] On the 928-instance DocCite-Syn OCR replication, stronger MinerU OCR improved supported recall but did not close the binary-accuracy gap to claim-level AtomCite: all nine family-by-OCR-condition paired gaps remained Holm-significant. The prose reports recall improvements of 6.6–15.1 percentage points; the table caption rounds this differently to 8–15 points.

Appendix B, OCR-engine robustness (full breakdown); Table A3 (n=928; doc-cluster bootstrap, B=10^4; Holm alpha=0.05 across nine comparisons) · Read the study

[c9] Human auditing overturned the initial natural-error ranking: among 608 instances disputed by claim-level AtomCite, 72.9% were supported label noise, 17.6% partially supported, and 9.5% annotator-mixed. With partial support counted as error, reported detection was 97.3–98.6%, and its misses were partial-support items. Both census waves ultimately yielded 1,909 validated errors, 450 supported claims, and 109 mixed items excluded from denominators; the undisputed region was sampled rather than fully human-adjudicated.

Appendix B, Natural-error audit trail; Second census wave; Appendix D, Figure A2 · Read the study

[c10] Main-matrix cells excluded 0–13 inconsistent or unparseable verdicts out of 928, although retry-exhaustion abstention was zero. Treating all exclusions as errors preserved ordering and 16 of 18 Holm-significant headline differences, but the Claude comparison against MinerU compute-matched OCR and the GPT comparison against the sequential chain lost significance.

Appendix B, Per-condition exclusions and the excluded-as-error recompute · Read the study

[c11] For natural-error correction, fix-page proposals landed within expanded gold sets in 50.8%, 58.7%, and 51.6% of Gemini, Claude, and GPT cases respectively. These are gold-set membership results, not fully human-scored repair accuracy: gold sets had limited recall and 60.5% audited equivalent-evidence precision, and item-level human evaluation of proposed targets remains future work.

Appendix B, Correction on natural errors · Read the study

[c12] Transfer evidence explicitly limits the framework to multi-claim long-form responses, with performance depending on evaluation unit. It lost 2.6 points on a 12,949-item pre-atomized pool, substantially underperformed direct judging on 300-item pilots of single-fact benchmarks, and changed from a 4.7-F1 loss to a 14.8-F1 gain when the same VeriGray data were evaluated per response rather than per sentence.

Appendix C, Boundary conditions · Read the study

[c13] The tested document domain was English business/legal document images. Other scripts, born-digital PDFs, and citation granularities other than pages were outside the study's scope.

Section 6, Limitations, item (iv) · Read the study

[c14] Primary scoring uses deterministic page localization as a proxy for support, not direct semantic entailment: a citation is correctly localized when its cited page belongs to the gold page set. The headline accuracy measure should explicitly retain this distinction.

Section 3, Two-layer ground truth · Read the study

[c15] Separate human calibration found a 7.2% localization-versus-semantic discrepancy in the injected strata, with reported interval [3.1, 15.9]. Passing calibration does not make the primary localization labels equivalent to semantic entailment judgments.

Section 5.6, Natural Errors, final paragraph · Read the study

[c16] On audited natural candidate errors, correction retained 88.0–94.7% of human-confirmed supported claims, demonstrating supported-content loss as well as error correction. These results concern the selected candidate-error pool, not deployment-wide citations; the supported subset comprised 450 claims.

Section 5.7, Correction; supported-subset denominator and candidate-pool scope in Table 4 caption · Read the study

[c17] In the experimental conditions ladder, the full AtomCite verifier (C) spent 2–3 times the layout-OCR condition's list-price cost. This is a measured experimental cost–accuracy comparison, not a customer deployment-cost estimate.

Section 6, What the results support · Read the study

[c18] In the experimental run ledger, default AtomCite averaged $0.024 and 2.8 calls per item, versus $0.009 and 1.3 calls for layout OCR: approximately 2.7 times the cost. These are uncached list-price averages, not customer deployment-cost estimates. Default AtomCite was measured on two cache-disabled repeats because its original run shared cached grounding calls.

Appendix B, Per-condition compute · Read the study

[c19] The reported cost comparison is specifically DocCite-Syn with the Gemini backbone. Call/token structure is identical across model families, but absolute costs differ by provider list price.

Appendix B, Table A2 caption and cache-disabled-repeat footnote · Read the study

[c20] Under conservative correction on the human-audited DocCite-Nat candidate-error pool, keep retention among 450 human-confirmed supported claims was 94.7% for Gemini, 94.7% for Claude and 88.0% for GPT. Thus supported-content loss was 5.3%, 5.3% and 12.0%, respectively. These conditional pool results are not deployment-wide citation retention rates.

Appendix B, Correction on natural errors; Appendix D, Figure A2 identifies the candidate-error pool construction · Read the study