The practical brief
The finding
Across five question-answering datasets and two models, training for sentence-level evidence support improved citation grounding and precision over prompting and adding citations afterward. [c2] [c3] [c5] [c6]
Why it matters
For a business choosing a customer-facing assistant, a citation label is not enough: the cited passage must support the claim. These results do not demonstrate greater brand visibility in public AI answers.
| Evaluation | Trained model, single pass | Trained model, iterative retrieval |
|---|---|---|
| NQ exact-match answer recall | 47.9% across 700 test queries | 51.0% across 700 test queries [c3] |
| ASQA exact-match answer recall | 35.7% across 948 test queries | 39.4% across 948 test queries [c3] |
| StrategyQA answer accuracy | 65.0% across 460 test queries | 64.6% across 460 test queries [c3] |
What to try
BLURSOR’s practical interpretation
Our interpretation: ask vendors to evaluate citation support and answer correctness separately on your own customer questions, and compare single-pass generation with additional retrieval before paying for either approach.
Study boundary
Citation scores relied on an automated evaluator, and the proprietary customer-support test had no reference answers for measuring correctness.
Buy evidence support, not just citation formatting
If you are choosing an AI assistant for customers, distinguish an answer that displays references from one whose references actually support its statements. The business risk is presenting unsupported information with the appearance of verification. The authors flag that risk, but do not measure whether citations persuade users.
This study offers evidence for evaluating an assistant’s credibility, not a marketing recipe for getting your business mentioned. It tests how models answer questions using retrieved text, rather than how public AI services select brands, rank websites or recommend suppliers.
The study trained models to recognize evidence gaps
The researchers tested text-bison-001 and llama-2-13b across five datasets, including general questions, ambiguous questions, multi-answer questions and proprietary customer-support questions. Their method, AGREE, fine-tunes a model—adds task-specific training—to generate sentence-level citations and identify unsupported statements.
Training examples were built automatically from questions without reference answers. A natural language inference model, an automated checker of whether a passage supports a statement, linked sentences to evidence. The best-grounded sampled response became a training example; it did not have to be fully supported.
An optional iterative stage then retrieved more passages using unsupported statements or the original question and generated another answer. The experiments allowed four model calls for this stage. Comparisons included prompting with citation examples and attaching citations after an answer was written.
Citation gains were clearer than universal accuracy gains
Across the five-dataset comparison, the trained approach improved citation grounding and precision over the baselines for both models. Grounding here means how well each sentence is supported by its cited passages; it is not a guarantee that the underlying corpus is correct.
Extra retrieval improved llama-2-13b exact-match answer recall from 47.9% to 51.0% on NQ’s 700 test questions and from 35.7% to 39.4% on ASQA’s 948 questions. This metric checks whether reference short answers appear in the response. However, StrategyQA accuracy fell from 65.0% to 64.6%, so additional retrieval was not uniformly beneficial.
The processing trade-off also matters. Single-pass AGREE averaged 1,210 language-model tokens per query, compared with 2,800 for the prompting baseline; iterative AGREE used 4,840. These are processed-token counts, not measured prices or response times. Fine-tuning also introduces a one-time adaptation cost.
Use separate checks before customer deployment
Our interpretation is to evaluate support, correctness and operational cost independently. Better citations can be useful without proving that an answer is complete or correct. The customer-support evaluation illustrates the distinction: its 580 questions had no ground-truth answers, so only citation quality was reported.
- Ask for a claim-by-claim review of whether cited passages support customer-facing answers, including partially supported claims.
- Compare extra retrieval against single-pass answers on your own questions; do not assume every accuracy metric will improve.
- Request actual pricing and response-time measurements rather than treating token counts as a deployment budget.
The evidence remains narrower than a deployment guarantee
The same type of automated inference model helped create training citations and evaluate grounding. The authors acknowledge that it may struggle to distinguish partial support. Human review therefore remains an important independent check, rather than something these scores replace.
The work primarily covers English information-seeking questions using two base models. It does not establish equivalent results for other languages, broader long-form writing, current-web discovery or business outcomes.
The source and evidence
Effective Large Language Model Adaptation for Improved Grounding and Citation Generation
Study limitations and disclosures
- Training supervision and citation evaluation depended on an automated natural language inference model rather than human citation judgments; partial support may be misclassified.
- The proprietary Enterprise dataset contained 580 customer-support queries without reference answers. Its results measure citation quality, not independently verified answer accuracy.
- The experiments primarily cover English information-seeking question answering with two base models, not public AI brand visibility, marketing outcomes or equivalent performance in other languages.
- Token-processing comparisons do not establish monetary savings or latency. Model adaptation also requires additional one-time training.
- Xi Ye was affiliated with The University of Texas at Austin and performed the work during an internship at Google Cloud AI. Ruoxi Sun, Sercan Ö. Arık and Tomas Pfister were affiliated with Google Cloud AI. The study tested Google's text-bison-001 alongside llama-2-13b; no separate funding statement is provided.
- The possibility that plausible but incorrect citations make unsupported claims more convincing is an author-identified risk, not a measured user-study result.
Evidence behind this briefing
[c1] AGREE trains a model on automatically selected, better-grounded responses and explicitly teaches it to identify unsupported statements; it does not require every training response to be fully supported. Appendix A specifies four sampled responses per query, five retrieved passages, a citation threshold above 0.7 and an unsupported-statement threshold of no passage scoring above 0.5.
Section 4.1, page 4; operational details in Appendix A, page 13 · Read the study
[c2] In the reported five-dataset comparison, AGREE improves citation grounding and precision over prompting and post-hoc citation baselines for both tested models. This is evidence about citation support within retrieved corpora, not evidence that a business will be mentioned more often in public AI answers.
Section 5.2, “Tuning is effective for superior grounding,” page 7; Table 2, page 6 · Read the study
[c3] With a four-LLM-call test-time budget, iterative retrieval improved llama-2-13b exact-match answer recall by 3.1 percentage points on NQ (47.9 to 51.0; 700 test queries) and 3.7 points on ASQA (35.7 to 39.4; 948). These are specific correctness gains, not universal improvements: Table 2 also reports StrategyQA accuracy declining from 65.0 to 64.6 for that model.
Section 5.2, “TTA improves both grounding and answer correctness,” page 7; Table 2 and footnote 5, page 6; Table 1, page 5 · Read the study
[c4] Average processed tokens per query were 1,210 LLM tokens for AGREE without iterative adaptation, versus 2,800 for the in-context citation baseline. Iterative AGREE used 4,840 LLM tokens. These are token-processing comparisons, not measured monetary costs or latency; post-hoc attribution additionally processed 3,520 tokens through a T5-11B NLI model.
Table 4, page 8 · Read the study
[c5] The 580-query Enterprise evaluation measures citation quality only: its proprietary customer-support dataset has no ground-truth answers. Consequently, its results should not be presented as independently measured answer accuracy.
Section 5.1, “Metrics,” page 7; proprietary-dataset disclosure in Section 5.1, page 6; Table 1, page 5 · Read the study
[c6] Both synthetic training citations and grounding evaluation depend on an NLI model rather than human citation judgments. The authors disclose that this model may not effectively distinguish partial support, limiting how confidently citation scores can be interpreted as trustworthy evidence.
Section 7, “Limitations and future work,” page 9 · Read the study
[c7] The study primarily covers English information-seeking question answering. It does not establish equivalent performance for other languages or broader long-form generation tasks.
Section 7, “Limitations and future work,” page 9; continuation on page 10 · Read the study
[c8] A citation can increase perceived credibility without actually supporting a statement. The authors identify this as a potential risk, not a measured user-persuasion result.
Section 7, “Limitations and future work,” page 10 · Read the study