The practical brief

The finding

Across 1,150 correct-answer attempts with separable answer and evidence layers, 201 lacked complete credited evidence: 17.48%. [c2] [c4] [c5] [c6]

Why it matters

A correct-looking financial answer is not necessarily ready to support a business decision or a published claim.

Three measures answer different business questions
MeasureReported result and scope
Best individual-criterion scoreClaude-Opus-5: 87.59% across 701 unweighted criteria on 123 financial tasks. [c3]
Best complete-task pass rateGPT-5.6-Sol: 71.54%, or 88 of 123 financial tasks fully satisfied; missing outputs counted as failures. [c3] [c7]
Correct answers missing complete evidenceAcross configurations: 17.48%, or 201 of 1,150 correct-answer attempts among 1,770 separable attempts. [c5]
Measured results from the shared financial-search harness. Partial credit, complete task success and evidence completeness have different denominators. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: review the evidence chain and calculation separately from the conclusion before reusing consequential AI-generated financial claims.

Study boundary

This financial-research benchmark does not measure brand discovery, citation frequency or the effects of website changes.

Check the evidence before publishing the number

Before using an AI-generated financial figure in a business plan, sales presentation or marketing claim, decide whether you can verify its supporting evidence—not just whether the number looks right. This study identifies a specific risk: answers can meet the benchmark’s correctness requirements while still missing required sources or other supporting evidence.

Across 1,770 attempts where researchers could score answers and evidence separately, 1,150 answers were correct. Of those, 201 lacked complete credited evidence, producing a 17.48% unsupported-correct rate. That percentage is conditional on answer correctness; it is not the overall error rate, nor does it mean those answers were false.

Evidence [c5]

The test separated finding facts from finishing research

FinFIRST V1 tested 15 model configurations on 123 expert-authored financial-research questions. Each used the same web-search, page-visit and Python calculation tools, reducing differences caused by proprietary search systems. Researchers reported one designated run per configuration.

The benchmark divided tasks into atomic criteria—small, independently scored requirements. These covered acquiring the right information, verifying sources, and forming the answer through calculation and reporting. Its source standard favored statutory, official and first-party evidence, while allowing suitable corroborated secondary sources.

Two scores answered different questions. An atomic score gave credit for individual requirements; a strict pass required every criterion in a task to be satisfied. Claude-Opus-5 led on atomic scoring at 87.59% across 701 criteria. GPT-5.6-Sol led on strict completion, fully satisfying 88 of 123 tasks, or 71.54%. These are different measures, not interchangeable accuracy rankings.

Evidence [c1] [c2] [c3]

Retrieval success did not guarantee a usable conclusion

Every evaluated model scored lower on computation and answer formation than on raw-information acquisition. The weighted capability-score gap ranged from 5.27 to 21.80 percentage points. Finding relevant information was therefore not the end of the job: calculation, synthesis and complete reporting remained bottlenecks.

For business owners, the realistic interpretation is to separate three checks: whether the source is appropriate, whether the extracted fact matches the question, and whether the conclusion follows. A citation alone cannot settle all three. Equally, missing credited evidence is not proof that the system never found a source: the judge assessed only what appeared in the final response.

For marketers seeking credibility in AI answers, the source hierarchy is informative but not a visibility experiment. It describes what this benchmark rewarded, not what makes commercial AI products retrieve or cite a business.

Evidence [c1] [c4] [c6]

Use traceability as a review standard, not a growth tactic

Our interpretation is to use the benchmark’s distinctions as a practical review checklist for consequential financial claims. This is an editorial and verification approach, not a proven method for improving AI discoverability.

  • Inspect the cited document and confirm the entity, reporting period, units and definition behind the figure.
  • Reproduce material calculations rather than accepting a plausible final percentage.
  • When publishing your own financial claims, make their source and calculation explicit so readers can check them; this study did not test whether doing so increases AI citations.

Evidence [c1] [c4] [c6]

The results do not establish commercial search performance

The study concerns financial research under a shared basic-tool environment, not consumer recommendations or marketing discovery. One run per configuration cannot establish repeated-run reliability. Five configurations had incomplete output coverage; missing outputs counted as failures in the fixed task denominator.

The authors’ affiliations include Inclusion AI, Peking University and CICC. CICC’s investment banking team supported construction, and the evaluation included a model from affiliated InclusionAI. These relationships warrant disclosure, without themselves invalidating the findings.

Evidence [c2] [c7] [c8]

The source and evidence

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

Wenqing Wang, Haitao Xiang, Xinyi Zhao, Mingming Yin, Ying Zhong, Zhaoxin Huan, Qiheng Zhou, Jin Zhu, Xiaolu Zhang, Shi Chang, Jun Zhou · 2026-09-21

Study limitations and disclosures

  • This benchmark evaluates financial research, not brand discovery, consumer preference, recommendation frequency or causal effects of website changes. Its shared search, page-visit and Python environment excludes additional professional financial tools and proprietary plugins, so the rankings should not be transferred directly to differently equipped commercial products.
  • The 123-task sample comprises 60.2% Chinese and 39.8% English questions. Companies are the largest research-object category, with mainland China and the United States contributing comparable geographic shares. This composition limits generalization to other languages, markets and business tasks.
  • Results represent one designated run per configuration at temperature 1.0, the generation-randomness setting. Repeated-run uncertainty estimates were not supplied. Five configurations had incomplete coverage: GPT-5.6-Sol produced 117 valid outputs and four others produced 122. Missing outputs counted as failed criteria across the fixed 123-task denominator; tool-efficiency statistics used recorded trajectories only.
  • The automated judge, GLM-5.1, scored final responses rather than hidden reasoning or retrieval trajectories. Validation against expert labels on 50 randomly sampled instances yielded item-level Cohen’s Kappa of 0.816, an agreement measure rather than proof of error-free grading. Evidence scores are not a complete audit of search behavior.
  • The unsupported-correct rate measures incomplete credited evidence among correct-answer attempts, not overall accuracy. Its analysis excluded five questions per model because their answer and evidence layers were not independently separable.
  • Benchmark construction used Claude-Opus-5, GPT-5.6-Sol and GLM-5.3 for stress testing; those models also appear in the evaluation. The authors state that experts controlled revisions and acceptance decisions.
  • Listed author affiliations include Ling Team, Inclusion AI; Peking University; and China International Capital Corporation Limited (CICC). The paper discloses professional construction support from CICC’s investment banking team and internship work at Ling Team, Inclusion AI. It also evaluates InclusionAI’s Ling-3.0-Flash-Fin, creating a material institutional and commercial relationship. The supplied paper does not specify a separate funding award or funding amount; construction support should not be represented as a quantified grant.

Evidence behind this briefing

[c1] The benchmark prioritizes statutory, official and first-party evidence, while accepting appropriate, corroborated secondary sources. This is an evaluation standard, not an experiment showing that particular publishing tactics increase AI visibility.

Section 2.1.3, Financial Source Map · Read the study

[c2] All 15 configurations were evaluated on 123 questions with the same ReAct-style protocol and search, visit and Python tools. Results represent one designated run per configuration at temperature 1.0, rather than repeated-run estimates.

Section 3.1, Evaluation Setup · Read the study

[c3] Partial-credit leadership differed from complete-task leadership: Claude-Opus-5 scored 87.59% on 701 unweighted atomic criteria and 87.61% on 12,300 weighted rubric points; GPT-5.6-Sol fully satisfied 88 of 123 tasks (71.54%). Missing outputs were scored as failures in the fixed denominator.

Section 4.1, Overall Performance; metric denominators in Section 3.3; coverage disclosure in Table 3 caption · Read the study

[c4] Every evaluated model scored lower on computation and answer formation than on raw-information acquisition. The weighted capability-score gap ranged from 5.27 to 21.80 percentage points; finding information did not ensure a correct, complete answer.

Section 4.1, Overall Performance · Read the study

[c5] Across 1,770 attempts with separable answer and evidence layers, 201 of 1,150 correct answers lacked complete credited evidence: a micro-averaged unsupported-correct rate of 17.48%. This measures evidence completeness conditional on answer correctness, not overall answer accuracy or user preference. Table 4 excludes five nonseparable questions per model from this analysis.

Section 4.2, Unsupported-Correct Answers; Table 4 · Read the study

[c6] The judge assesses only evidence presented in final responses, not hidden reasoning or retrieval trajectories. Source-verification scores therefore should not be described as a complete audit of agent search behavior.

Appendix C, Judge Prompt Template · Read the study

[c7] Result coverage was incomplete for five configurations, especially GPT-5.6-Sol with 117 valid outputs. Missing outputs counted as failed criteria across the fixed 123-task denominator, while tool-efficiency statistics used only recorded trajectories.

Table 3 caption · Read the study

[c8] The construction process received professional support from CICC's investment banking team. The supplied affiliations also identify Ling Team, Inclusion AI; the evaluation includes InclusionAI's Ling-3.0-Flash-Fin model. These are material institutional relationships, not evidence that the results are invalid.

Section 1, Introduction; opening affiliations; Appendix B, Table 6 · Read the study