The practical brief
The finding
In a GPT-4.1 earnings-call benchmark, citation precision averaged 0.9721, while numeric claim precision averaged 0.6296 against automatically constructed references. [c1] [c3] [c5]
Why it matters
A cited competitor or company summary can look well supported while still containing unmatched figures or omitting relevant evidence.
| Benchmark measure | Mean score and scope |
|---|---|
| Citation precision | 0.9721 (SE 0.0032), mean score across 100 transcript documents [c3] |
| Numeric claim precision | 0.6296 (SE 0.0083), mean score across 100 transcript documents [c3] |
| Citation recall | 0.7139 (SE 0.0087), mean score across 100 transcript documents [c3] |
What to try
BLURSOR’s practical interpretation
Our interpretation: check consequential figures and their context separately from whether a citation points to the expected passage.
Study boundary
This tests numeric alignment in public earnings-call transcripts, not complete factual accuracy, brand discoverability or marketing outcomes.
Do not let citations substitute for checking figures
If you use AI to summarize competitors’ earnings calls or prepare commentary about your own market, the decision is whether a cited answer is ready to reuse. This study suggests that citation quality and numeric answer quality need separate checks. A reference can match the expected source passage without making the generated analysis fully correct.
The researchers evaluated GPT-4.1 on fourth-quarter 2024 earnings-call transcripts for 100 S&P 500 companies, substituting The Cigna Group for Berkshire Hathaway. They tested six citation-generation configurations, then used a sentence-identifier configuration for further evaluation. These were different ways of using one model—not a comparison across AI providers.
What the benchmark actually measures
The study’s numeric evidence evaluation checks whether numbers match after normalizing their type, value and unit. Groundedness here means numeric alignment with the supplied transcript or citation; it is not an expert judgment that the whole sentence follows from the evidence.
For correctness testing, the authors automatically constructed questions, reference answers and reference citations from salient figures. Citation precision measures how much of the generated citation evidence matches those references. Numeric claim precision similarly measures matching generated figures. Recall measures how much reference content the answer recovers.
The selected configuration achieved citation precision of 0.9721, but numeric claim precision of 0.6296. Citation recall was 0.7139, indicating incomplete reference coverage. These are average scores, not the proportions of answers that were entirely correct.
Missing evidence creates a second risk
The researchers also removed the reference passages containing target figures while leaving each question and the remaining transcript unchanged. The desired behavior was to say that the supplied document could not answer the question.
Instead, the average document-level hallucination rate was 0.4834, with a standard error of 0.0149. In this experiment, “hallucination” specifically meant producing claims containing salient numbers or supporting citations despite the simulated evidence gap. It is not a general rate for financial AI use.
For a business reader, the realistic interpretation is that fluent, cited output should not be mistaken for evidence sufficiency. An answer may continue even when the material needed for that question has been deliberately removed.
Evidence [c4]
Use the findings as a review checklist, not a visibility tactic
Our interpretation is to review consequential AI-generated financial statements at two levels: whether the cited passage is relevant, and whether the claimed number means what the answer says it means. That is a practical precaution, not a workflow improvement tested by this study.
- Check the figure, unit, company, financial concept and reporting period against the passage.
- Look for missing evidence before treating a summary as complete.
- Keep unsupported conclusions out of publishable commentary, or clearly label them as unresolved.
What this study cannot tell your business
Identical numbers can count as matches even when they describe different entities, concepts or periods. The metric does not directly test nonnumeric claims, and it was not validated against human or expert judgments. Cases without numeric content receive a score of 1.
The automatically generated references also limit the meaning of “correctness.” These findings do not establish performance across model families, prove that a review process works, or show how to earn more public AI citations. They concern a financial-document benchmark, not customer-facing discoverability.
The source and evidence
Study limitations and disclosures
- All six configurations used GPT-4.1, which also generated dataset questions. References were automatically constructed rather than established by financial experts; numeric alignment was not validated against human judgments.
- Dataset targets were restricted to salient figures cited in an earlier benchmark. Reported standard errors describe document-level variability; the authors report no statistical analysis establishing significance.
- The authors are affiliated with American Express and explicitly state that the paper does not represent the company’s or its affiliates’ views, policies or practices. The verified disclosure does not establish a separate funding arrangement.
Evidence behind this briefing
[c1] Numeric evidence evaluation compares normalized numeric type, value and unit rather than evaluating full semantic support. Measurement-bearing figures are extracted using structured output; dates and identifier-like numbers are excluded, ranges must match as ranges, and cases with no numeric content receive a score of 1. Correctness additionally compares generated figures and citation identifiers against automated references, averaging QA-instance scores within documents and reporting means and standard errors across documents.
Sections 2.2 and 2.4; Appendix D · Read the study
[c2] Across six zero-shot GPT-4.1 citation-generation configurations tested on five thematic questions over full transcripts, numeric claim–context groundedness ranged from 0.970 to 0.982, citation–context groundedness from 0.989 to 0.995, and claim–citation groundedness from 0.907 to 0.959. LLMR-text had the highest claim–citation score; LLMR-index was selected for subsequent evaluation because its sentence identifiers support reference matching. These are numeric-alignment scores, not human-assessed factual accuracy; Table 1 supplies no uncertainty estimates.
Sections 3.1, 3.2.1 and 3.2.2; Table 1 · Read the study
[c3] On automatically constructed ECTs-100 references, the selected LLMR-index strategy achieved citation precision 0.9721 (SE 0.0032), recall 0.7139 (0.0087), and F1 0.8204 (0.0063). Numeric claim precision was 0.6296 (0.0083), recall 0.6658 (0.0079), and F1 0.6468 (0.0080). Scores are QA-instance averages within each document, then means and SEs across the 100 documents, not percentages of fully correct answers. Dataset construction restricts targets to salient figures appearing in at least one citation in the earlier groundedness benchmark.
Sections 3.3.1–3.3.2; Section 2.4.2; Appendix D.2 · Read the study
[c4] In an evidence-removal stress test on ECTs-100, the average document-level hallucination rate was 0.4834 (SE 0.0149). The study defines hallucination here as generating claims with salient numbers or supporting citations after target-figure reference citations have been removed. This measures failure to abstain under the study's simulated insufficiency condition, not a general real-world hallucination rate.
Section 3.4, Conscious incompetence; Section 2.4.2 · Read the study
[c5] The authors explicitly limit NEE to a numeric-alignment proxy: equal figures can falsely match across entities, concepts or reporting periods; nonnumeric support is untested; no human/expert validation is supplied. All six configurations use one underlying model, GPT-4.1, which also generates dataset questions, so references are automatically constructed rather than expert-established.
Appendix A, Limitations · Read the study
[c6] The paper is an academic contribution by authors affiliated with American Express, not evidence of American Express's business views, policies or practices.
Section 5, Disclaimer; author affiliations · Read the study
[c7] The authors disclose that they perform no statistical analysis; their assertion that a substantial number of test ECTs improves stability is not a reported significance test or uncertainty estimate.
Checklist item 7, Experiment statistical significance — Justification · Read the study
[c8] This is an evaluation study without model training. The authors state that they specify testing datasets, task definitions, default settings, and evaluation metrics.
Checklist item 6, Experimental setting/details — Justification · Read the study
[c9] The authors report detailed benchmark-construction documentation and GitHub availability as support for reproducibility; this chunk does not independently establish successful replication.
Checklist item 4, Experimental result reproducibility — Justification · Read the study
[c10] The evaluation concerns proprietary state-of-the-art models accessed through APIs, with downstream experiments described as lightweight laptop work.
Checklist item 8, Experiments compute resources — Justification · Read the study
[c11] The authors disclose exclusive use of public datasets and API-based model access, without modifying or bypassing built-in LLM safety mechanisms.
Checklist item 11, Safeguards — Justification · Read the study
[c12] The LLM-usage declaration states that LLMs were used only for grammar editing. This declaration should be distinguished from the separately disclosed evaluation of LLMs through APIs.
Checklist item 16, Declaration of LLM usage — Justification · Read the study