The practical brief

The finding

Retrieving content does not guarantee that an AI answer will reference it, and a reference does not establish answer correctness. [c1] [c2] [c3] [c8] [c25]

Why it matters

Business owners evaluating AI visibility services need to distinguish source mentions from accurate, evidence-supported representation.

Three checks a citation count cannot replace
CheckWhat it establishes—and leaves open
Citation alignmentReferences match annotated supporting sources; this alone does not establish support for the generated claims. [c23]
Claim supportEvidence supports a statement; this does not establish that the model actually used that evidence during generation. [c20]
Answer correctnessRequires separate assessment; automated scores provide partial signals, and model-based judging requires alignment with human judgments. [c3] [c25]
Qualitative interpretation of the survey’s evaluation distinctions; not measured business outcomes. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: assess citation presence, claim support, and correctness separately, and ask how any automated correctness scores are validated against human judgments.

Study boundary

This is a literature survey with a February 2025 search cutoff, not a marketing experiment. Automated correctness metrics provide partial signals, not comprehensive factual assessment.

Decide what a citation report actually proves

Before paying for an AI visibility service, decide whether you want evidence of exposure or evidence that answers describe your business correctly. Seeing your website in an answer’s references establishes a visible reference—not necessarily accurate representation.

The survey’s central distinction is practical: retrieving evidence, attaching references, and producing correct statements are separate capabilities. Retrieval-augmented generation, or RAG, gives a language model external documents to work with. The authors state that standard RAG still needs additional mechanisms to explicitly reference those documents.

The business risk is treating a citation dashboard as a complete quality report. The paper supports separating these checks; it does not measure how often commercial dashboards conflate them.

Evidence [c2] [c3]

What the researchers studied

Tobias Schreieder, Tim Schopf, and Michael Färber conducted a systematic mapping study: a structured search and classification of research, rather than a new product experiment. Their February 2025 search covered nine literature databases. After deduplication, they manually screened 805 publications and included 134.

Their public dataset contains those annotated publications, 300 extracted evaluation metrics, and 231 datasets. These counts describe research resources—not commercial answer accuracy or opportunities to win citations.

The authors recommend evaluating correctness separately from attribution or citation. Attribution asks whether evidence supports a generated claim. Citation evaluation can examine whether references match annotated supporting sources, but that match alone does not establish support for the generated statement.

Evidence [c1] [c3] [c7] [c8] [c23]

Ask for answers, not just aggregate scores

Our interpretation is to request an inspectable sample of answers alongside aggregate citation reporting. Keep three questions distinct: was your source referenced, does it support the attached statement, and is the statement correct?

Even support is not the whole story. The survey distinguishes a supporting citation from a faithful one: a source can support a statement without having actually been used to generate it. A reference is not a transparent record of the model’s reasoning.

A separate automated correctness score is not definitive either. The survey describes a trade-off between scalability and semantic coverage: automated checks can process more text, but capture only partial signals and can miss nuanced factual errors. LLM-as-a-judge—using a language model to evaluate another answer—extends coverage but requires careful alignment with human judgments.

The review also reports user studies where citations increased perceived trust. That concerns presentation, not verified accuracy or a measured benefit for your business; the supplied evidence does not give full experimental conditions, effect sizes, or uncertainty.

  • Ask what each metric measures: reference presence, source support, or correctness.
  • Ask how correctness scores are validated, how model-based judgments compare with human reviews, and which errors the checks may miss.
  • Treat content changes as hypotheses to test, not proven ways to obtain more citations.

Evidence [c3] [c16] [c17] [c20] [c23] [c25]

What this survey cannot tell your business

The paper establishes no causal tactic for improving discoverability, recommendations, or citation frequency in deployed AI products. Its search cutoff is February 2025 despite its 2026 publication metadata. English-language and public-source restrictions may underrepresent regional and proprietary research.

Coverage and evaluation are uneven. The datasets predominantly concern question answering, with frequently reused resources concentrated in Wikipedia, scientific, and news domains. Only two of 19 evaluation frameworks and two of 11 benchmarks were reused across multiple studies. Automated correctness metrics remain indicative rather than comprehensive.

Researcher judgment shaped screening and classification. An expanded search in two databases found five additional relevant studies, excluded from the main corpus because they did not change the taxonomy. The authors disclose German and Saxon public funding, Software Campus support, and a DAAD scholarship, plus AI assistance for editing, formatting, and coding, with author review and no AI contribution to scientific conclusions.

Use the survey to sharpen procurement questions—not to forecast citation gains, customer trust, or returns from a content rewrite.

Evidence [c5] [c6] [c8] [c12] [c13] [c18] [c22] [c25]

The source and evidence

Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models

Tobias Schreieder, Tim Schopf, Michael Färber · 2026

Study limitations and disclosures

  • The February 2025 search included only five papers from 2025; the publication date should not be treated as the literature-search cutoff.
  • The authors acknowledge that a single primary search string may have omitted relevant studies. An expanded-term sensitivity analysis identified a reported 4% additional proportion of relevant studies; methodological details are deferred to Appendix B, which is unavailable in this chunk.
  • Inclusion screening and categorization involved researcher judgment and possible bias. The authors report explicit inclusion criteria, consensus discussions, inter-annotator agreement measurement, and iterative schema refinement, but this chunk does not provide the agreement results.
  • Automated correctness metrics provide indicative, partial signals rather than comprehensive factual assessment. The authors describe a scalability–semantic-coverage trade-off and caution that LLM-as-a-judge methods require alignment with human judgments.
  • Funding disclosures name German BMFTR and Saxon ministry support for ScaDS.AI, BMFTR Software Campus support for Tobias Schreieder under project 16|S23070, and a DAAD scholarship for Tim Schopf.
  • The survey does not establish causal tactics for improving a business's discoverability, citation frequency, or credibility in commercial AI answers.
  • The corpus reflects a February 2025 search with English-language and topic-specific eligibility restrictions, not unrestricted coverage of evidence-grounded generation research.
  • Search recall was assessed against an initial literature list; sensitivity testing covered only two of the nine databases and found five relevant studies omitted from the main corpus.
  • The reported perfect agreement concerns a random sample of 50 inclusion-screening decisions, not all annotations.
  • The chunk does not supply sample sizes, effect sizes, uncertainty intervals, or full experimental conditions for the secondary-reported Pride, Do, and Ding results; do not generalize them to all models, users, or current commercial systems.
  • Annotation agreement was assessed on a random subset of 30 publications. Label-level Krippendorff’sα was calculated only for labels appearing at least five times and macro-averaged within dimensions. No parametric-attribution studies were present in that subset, so agreement for that dimension could not be computed.
  • Studies can receive multiple labels; contribution, task, and evidence-level totals can exceed the 134-paper denominator.
  • The surveyed evidence is predominantly textual and concentrated in encyclopedic sources, scientific literature, and general web content. Specialized or regulated domains are underrepresented; non-textual modalities have too few studies for meaningful task-specific differentiation.
  • Temporal coverage includes relevant Q1 2025 papers available up to the February inclusion deadline, rather than the complete quarter.
  • Appendix E.3 describes selected notable metrics rather than all 300 extracted metrics; the complete categorized table is disclosed as available in the repository.
  • Table 7 orders papers by the earliest accessible version, including preprints; temporal adoption counts should not be interpreted as dates of final peer-reviewed publication.
  • Table 12 is only a sample of the 231 datasets. Its caption discloses that LLM-Rubric and STRoGeNS each contain multiple datasets combined for simplicity.
  • The reported unfaithful-citation percentage is a secondary report without the underlying test-system details or denominator in this chunk.
  • Assess answer correctness separately, but do not treat a separate automated correctness score as definitive. As an editorial implication of the survey’s limitation, ask vendors how their correctness scores are validated, including how LLM-as-a-judge assessments are aligned with human judgments and which factual errors the metrics may miss.
  • Editorial implication of the supported metric limitations: ask vendors what their correctness scores measure and how they are validated; do not treat a separate automated score as a comprehensive factual assessment. This is advice derived from the limitations, not a measured result or a verbatim survey recommendation.

Evidence behind this briefing

[c1] The survey used a PRISMA systematic mapping study: 805 deduplicated papers were manually screened, with 134 judged relevant. These are literature-sample counts, not model-performance measurements.

Page 2, Section 3, Evidence-based Text Generation · Read the study

[c2] Retrieval alone does not guarantee explicit source attribution: the survey states that standard RAG needs additional mechanisms to reference retrieved documents.

Page 3, Section 3.1.2, Non-Parametric Attribution · Read the study

[c3] The authors recommend evaluating correctness separately from attribution or citation. Attribution versus citation evaluation depends on evidence availability; contextual dimensions depend on the task and system design. This is evaluation guidance, not evidence that cited answers are accurate.

Page 8, Section 5.3, Evaluation Guidelines · Read the study

[c4] Text was the citation modality in 96% of surveyed studies. This describes the research literature, not the percentage of sources cited by deployed AI products; multimodal evidence remains underexplored.

Page 4, Section 3.2.1, Citation Modality · Read the study

[c5] English-language and publicly accessible source restrictions may underrepresent non-English, regional, non-public, and industry-internal research. Including arXiv did not resolve the limited visibility of proprietary research.

Page 10, Limitations · Read the study

[c6] The authors disclose AI assistance for language editing, minor formatting, and coding, while stating that the tools did not contribute intellectual content or scientific conclusions and that authors reviewed all content.

Page 10, Acknowledgments · Read the study

[c7] The authors disclose a publicly available BSD 3-Clause dataset containing 134 annotated publications, 300 extracted evaluation metrics, and 231 datasets.

Page 22, Appendix A: Data Availability · Read the study

[c8] The systematic mapping study used a title-and-abstract keyword search across nine literature databases. The February 2025 search returned 1,040 records, reduced to 805 unique publications after deduplication.

Page 23, Appendix B.1: Literature Search and Inclusion Criteria · Read the study

[c9] Eligibility was restricted to English, electronically accessible full texts about LLM natural-language generation that deliberately incorporates evidence-source references. The authors report that accessibility excluded no otherwise eligible publications.

Page 23, Appendix B.1: inclusion criteria · Read the study

[c10] Search design prioritized feasible manual screening rather than LLM-based screening. Stemming the citation, attribution, and quote keywords increased retrieved papers by 1,207%, exceeding 10,000 publications, with negligible recall improvement relative to the authors' initial literature set.

Page 23, Appendix B.1: search-string selection · Read the study

[c11] For inclusion screening, two annotators independently labeled a random sample of 50 studies before discussion; reported Krippendorff’s alpha was 1.0. This is screening agreement, not generation accuracy or agreement across the entire corpus.

Page 23, Appendix B.1: Inter-Annotator Agreement · Read the study

[c12] Sensitivity testing was limited to ACL Anthology and arXiv, covering 84% of included studies. It retrieved 126 studies and screened 103 unique additional papers, finding five relevant studies; the authors excluded those five from the main corpus because they did not change the taxonomy structure.

Page 23, Appendix B.1: Sensitivity Analysis · Read the study

[c13] The survey categorized 134 studies using annotations split between the first two authors, regular calibration, and multi-label assignments where needed.

Appendix B.2, Categorization Scheme and Method · Read the study

[c14] The survey reports that Pride et al. (2023) found factual citation correctness of 22% for GPT-3.5 and 20% for GPT-4. These are historical, secondary-reported citation-correctness results, not answer-accuracy rates or current-system benchmarks.

Appendix D.1.1, Parametric Attribution, Pure LLMs; page 25 · Read the study

[c15] Non-parametric attribution depends on accurate available evidence and retrieval quality, and does not itself expose the model’s reasoning.

Appendix D.1.2, Non-Parametric Attribution, Task-specific Analysis; page 28 · Read the study

[c16] The survey reports a Do et al. (2024) user study in which cited responses were perceived as significantly more trustworthy than uncited responses, with no significant trust difference between in-line citations and highlight gradients. This concerns perceived trust, not verified accuracy.

Appendix D.2.3, Citation Style; page 31 · Read the study

[c17] The survey reports that Ding et al. (2025) found greater perceived trust when citations were present, but no additional trust gains from adding more than one citation.

Appendix D.2.5, Citation Frequency; page 33 · Read the study

[c18] Among the evaluation resources identified by the survey, only two of 19 frameworks and two of 11 benchmarks were reused across multiple studies, limiting evaluation standardization.

Appendix E, Evaluation Resources; page 34 · Read the study

[c19] Across tasks in the annotated survey, attribution, citation and correctness account for 64%–81% of evaluation instances. Citation evaluation is conditional on annotated ground-truth evidence; fact verification instead relies primarily on attribution and correctness.

Appendix E.2, Task-specific Analysis, p. 34 (30989); qualification continues on p. 35 (30990) · Read the study

[c20] The survey distinguishes source support from faithfulness: a citation may support a claim without having actually been used during generation. It reports Wallat et al.'s 'up to 57%' unfaithful-citation result, but this chunk does not provide that study's systems, sample denominator or uncertainty, so the percentage should not be generalized to all RAG citations.

Appendix E.3, Attribution, p. 36 (30991) · Read the study

[c21] ALCE is the most widely reused complete evaluation framework in this survey, appearing in 12 studies and covering attribution, correctness and linguistic quality. Studies using only individual ALCE metrics are excluded from those 12 instances.

Appendix E.4, Frameworks, p. 36 (30991); Table 10, p. 44 (30999) · Read the study

[c22] The 231 extracted datasets include both training and evaluation resources and are highly task-specific. More than 64% concern question answering, versus 9% grounded text generation and 6% summarization; frequent datasets primarily cover Wikipedia, scientific and news domains, limiting direct extrapolation to other domains.

Appendix E.6, Datasets, p. 37 (30992) · Read the study

[c23] Citation Retrieval measures alignment with supporting-source annotations in question answering and summarization. It requires ground-truth citations, depends on the oracle set's completeness and does not establish whether cited sources support generated claims.

Table 9, Citation Retrieval row, p. 43 (30998) · Read the study

[c24] FActScore evaluates claim-level attribution for long-form question answering, grounded text generation and fact verification. Its listed requirements include a pretrained NLI model and access to generation evidence; it measures factual precision rather than recall and typically requires human annotation for atomic-fact decomposition.

Table 9, FActScore row, p. 43 (30998) · Read the study

[c25] For correctness assessment in long-form text generation, the survey identifies a scalability–semantic-coverage trade-off: automated metrics capture partial, indicative signals rather than comprehensive factual accuracy, and LLM-as-a-judge assessments require careful alignment with human judgments.

Page 9 (printed page 30964), Section 5 Evaluation, Takeaways · Read the study

[c26] The survey qualifies Exact Match as a correctness metric for short factoid answers: it measures exact text matching rather than semantic correctness and is not applicable to long-form generation.

Page 43, Table 9, correctness (COR), Exact Match row · Read the study

[c27] For question answering and summarization, Claim Recall provides a partial correctness signal: it evaluates recall rather than precision, requires a pretrained NLI model, and depends on claim-decomposition accuracy.

Page 43, Table 9, correctness (COR), Claim Recall row · Read the study