The practical brief

The finding

Search improved established-fact accuracy across four systems, but accuracy on retrieval-required current questions ranged from 60.8% to 71.4%, and self-reported confidence became less well calibrated. [c1] [c2] [c4] [c6] [c13] [c14] [c15]

Why it matters

Enabling search does not ensure an AI system will use it when needed—or find the right evidence when it does.

Current-fact accuracy varied by retrieval workflow
ModelInternal web searchExternal retrieval workflow
GPT-5-mini68.4% accuracy ±5.4 percentage points; 288 dynamic queries87.8% accuracy ±6.1 percentage points; 288 dynamic queries [c2] [c3] [c13] [c14]
Claude Haiku 4.560.8% accuracy ±5.6 percentage points; 288 dynamic queries85.3% accuracy ±6.5 percentage points; 288 dynamic queries [c2] [c3] [c13] [c14]
GPT-571.4% accuracy ±5.3 percentage points; 288 dynamic queries90.9% accuracy ±7.1 percentage points; 288 dynamic queries [c2] [c3] [c13] [c14]
Claude Sonnet 4.666.2% accuracy ±5.1 percentage points; 288 dynamic queries88.1% accuracy ±6.7 percentage points; 288 dynamic queries [c2] [c3] [c13] [c14]
Measured accuracy on 288 dynamic queries per model during 15–20 August 2026; reference answers were verified as of 23 August. Provider-controlled live search can change, so these are not future-performance guarantees. Values reproduce Table 3’s uncertainty, whose notation is undefined there; do not interpret external ± values as 95% confidence intervals. Table 2 separately labels internal accuracy uncertainty as 95% confidence intervals. External retrieval summarized two Google Custom Search results; the comparison changes the whole workflow. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: retain source checks for consequential claims and evaluate complete retrieval workflows on business-relevant questions.

Study boundary

This dated, English-language factual-lookup benchmark did not measure brand visibility, recommendations, customer decisions or marketing returns.

Keep a checking step for consequential facts

If your business uses AI for research or customer-facing content, do not treat a search-enabled answer as a verified fact. The decision supported here is whether to retain factual checks—not whether to abandon AI search.

Search improved established-fact accuracy in every tested system. Yet on questions requiring current information, accuracy ranged from 60.8% to 71.4%. The strongest result exceeds the abstract’s claim that accuracy remained below 70%; this briefing follows the results table.

For marketers, publishing correct information and having an AI system retrieve and use it correctly are separate problems. This study examined answer accuracy, not whether businesses gained mentions or recommendations.

Evidence [c1] [c2] [c15]

Four systems, simple lookups, one live-web snapshot

The black-box evaluation—testing outputs without access to model internals—covered GPT-5-mini, GPT-5, Claude Haiku 4.5 and Claude Sonnet 4.6. It used 783 static questions about established facts and 288 dynamic questions requiring information beyond the experimental knowledge cutoff.

Models autonomously decided whether to search, with at most two calls per question. Questions were designed so answers could be found in the first one or two standard search results. Correctness meant normalized exact matching against accepted answer variants, not a judgment of overall usefulness.

Experiments ran 15–20 August 2026; reference answers were verified as of 23 August 2026, after that window. These results are a live-web snapshot: provider-controlled retrieval can change over time.

On static questions, search-enabled accuracy ranged from 74.7% to 90.7%, versus 41.4% to 75.5% without search. Static and dynamic results address different information needs and should not be directly compared.

Evidence [c1] [c2] [c8] [c9] [c13] [c14]

Search can fail before evidence reaches the answer

On the retrieval-required dynamic set, search invocation ranged from 87.5% to 93.9%. Table 4 reports zero accuracy among no-search cases for all four systems. That is conditional accuracy within those subgroups, not across all 288 questions: making search available did not ensure its use.

An external Google Custom Search workflow performed better for every model. It summarized the top two results’ titles and snippets before answering. This compares complete configurations, not search engines in isolation.

Manual analysis mainly identified query formulation and source quality problems. It examined only wrong answers where search was invoked, excluding missed retrieval. Inconsistent sample counts and category percentages mean its conclusions are best treated qualitatively.

Confidence calibration—how closely stated certainty tracks correctness—also worsened with search. Confidence was a number supplied by the model, not a measurement of internal uncertainty.

  • Our interpretation: check underlying sources for consequential claims, rather than accepting confident wording.
  • Compare complete retrieval workflows on representative business questions before relying on them.
  • Do not infer an AI-visibility tactic from this accuracy benchmark.

Evidence [c3] [c4] [c5] [c15] [c16] [c17] [c18]

What this cannot tell your business

The benchmark does not establish citation frequency, website discoverability, sales impact or dependable performance on multilingual research and complex comparisons. Exact-match scoring measures factual agreement, not briefing completeness.

The two-call limit narrows its scope. Confidence-prompt sensitivity was checked on only 100 static questions in two models, not across arbitrary prompts.

The realistic takeaway is that search improved factual accuracy while leaving meaningful errors and inflated certainty. Source checking and workflow evaluation are our suggested safeguards, not proven marketing tactics.

Evidence [c1] [c4] [c6] [c7] [c9] [c12]

The source and evidence

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

Sahil Kale · 2025-11-24

Study limitations and disclosures

  • Four commercial model–search configurations were tested on simple English factual lookups with at most two search calls. Findings do not establish multilingual, complex browsing or marketing performance.
  • Experiments ran 15–20 August 2026, while reference answers were verified as of 23 August, after the experimental window. Provider-controlled live retrieval can change; this snapshot is not a temporally repeatable performance guarantee.
  • The external comparison changes the complete retrieval workflow, so its advantage cannot be attributed solely to the search engine. Table 3 does not define its ± notation; its external uncertainty values should not be called 95% confidence intervals. Table 2 separately labels accuracy uncertainty as 95% confidence intervals.
  • Failure-analysis sample counts and category percentages are internally inconsistent. Its categories are reported qualitatively and exclude failures to invoke search. Table 4 supplies neither subgroup counts nor uncertainty intervals for no-call accuracy.
  • Confidence was self-reported. The confidence-wording sensitivity check covered only 100 static questions and two models, not broad prompt robustness.

Evidence behind this briefing

[c1] The black-box API evaluation tested GPT-5-mini, GPT-5, Claude Haiku 4.5 and Claude Sonnet 4.6 on 783 static factual queries and 288 retrieval-required dynamic queries. Temperature was 0, retrieval was autonomous, and at most two search calls were allowed. Queries were deliberately simple enough for answers to be discoverable within the top one or two standard search results; correctness used exact matching against accepted answer variants.

Section 3.3, Experimental Setup; Table 1; Sections 3.1.2 and 3.2.2 · Read the study

[c2] On 783 static queries, enabling internal search increased answer accuracy for all four systems: GPT-5-mini from 52.3% to 84.6%, Haiku from 41.4% to 74.7%, GPT-5 from 75.5% to 90.7%, and Sonnet from 73.1% to 86.8%. On the separate 288-query dynamic split, accuracy ranged from 60.8% to 71.4%. Table 2 labels its reported uncertainty as 95% confidence intervals. These are end-to-end model–provider-search results; the two splits are not directly comparable.

Table 2; Sections 3.2.1 and 4.1 · Read the study

[c3] On the dynamic split, an external Google Custom Search pipeline outperformed internal search for each tested model: reported accuracy rose from 68.4% to 87.8% for GPT-5-mini, 60.8% to 85.3% for Haiku, 71.4% to 90.9% for GPT-5, and 66.2% to 88.1% for Sonnet. The external pipeline used the top two results’ titles and snippets, summarised by the same model before answering. This is a comparison of different end-to-end configurations, not an isolated test of search-engine quality; Table 3 reports uncertainty without explicitly defining it in its caption.

Section 4.1; Table 3 · Read the study

[c4] Although static-split accuracy improved with search, reported confidence calibration worsened across all four systems. Table 9 reports ECE increases from 0.10 to 0.27 for GPT-5-mini, 0.14 to 0.26 for Haiku, 0.10 to 0.17 for GPT-5, and 0.08 to 0.21 for Sonnet. Confidence was verbalised on a [0,1] scale, not measured from internal uncertainty; the reported prompt-sensitivity check covered only 100 randomly selected queries and the two frontier models.

Section 4.3; Table 9 · Read the study

[c5] The authors’ manual error analysis identifies query formulation and source quality as the main failure categories, rather than evidence integration. Its stated scope is cases where search was invoked but the final answer was wrong, with two graduate-level annotators, claimed random samples of 150 static and 75 dynamic errors per model, and Cohen’s kappa of 0.79. These sample claims and some category percentages have internal consistency problems, so the category-level conclusion should not be presented as a precise population estimate.

Section 4.5; Table 11 · Read the study

[c6] The benchmark is English-only, limits retrieval to two calls, and excludes multi-hop, compositional and adversarial retrieval tasks. Its confidence analysis uses verbalised confidence only. Consequently, these findings cannot be assumed to generalise to multilingual search or more complex browsing workflows.

Section 7, Limitations and Future Work · Read the study

[c7] The static split deliberately targets temporally anchored, unambiguous factual questions answerable through a single lookup, excluding multi-step reasoning and extensive synthesis. It therefore does not represent all web-search information needs.

Appendix A, A.1 Static Split · Read the study

[c8] Dynamic questions were manually checked against sources including Wikipedia revision histories and reputable news. Inclusion required a reliably established answer change after the experimental knowledge cutoff; this filtering produced 288 queries.

Appendix A, A.2.3 Post-cutoff Validation and Answer Verification · Read the study

[c9] Answer correctness used normalized exact matching against canonical answers or validated semantic variants, without manual correction of model outputs. Self-reported confidence was used only for calibration and retrieval-selectivity analyses, not correctness scoring.

Appendix B Evaluation Protocol · Read the study

[c10] The failure taxonomy applies specifically to queries where a model invoked web search but its final answer failed reference-answer matching. Such cases were assigned one mutually exclusive label: query formulation, source quality, or evidence integration failure.

Appendix C Failure Taxonomy Annotation · Read the study

[c11] Failure-label agreement was assessed by two independent annotators on 150 random incorrect static-split samples per model and 75 dynamic-split samples each. Cohen’s kappa was 0.79; disagreements were resolved by consensus for the reported analysis. This measures annotation consistency, not model accuracy.

Appendix C, C.0.2 Inter-Annotator Agreement · Read the study

[c12] For 100 static-split queries, confidence scores elicited by two semantically equivalent prompts had Pearson correlations of 0.88 without search and 0.86 with search for Claude Sonnet 4.6, and 0.85 and 0.84 respectively for GPT-5. This is confidence-wording stability in two models, not evidence of answer accuracy or robustness across arbitrary prompts.

Table 14; Appendix D, D.1 · Read the study

[c13] The dynamic split contained 288 retrieval-required queries. Experiments ran 15–20 August 2026, while reference answers were verified as of 23 August 2026, after the experimental window. Results represent a versioned, time-sensitive snapshot.

Section 3.2.1, paragraph beginning 'Following this process' · Read the study

[c14] Provider-controlled live search can change over time, so the reported snapshot is not a guarantee of temporally repeatable retrieval or future performance; this qualification also applies to visuals presenting the results.

Section 7, 'Limitations of live-web evaluation' · Read the study

[c15] On the dynamic split, where every query required retrieval, search invocation ranged from 87.5% to 93.9%. Enabling search therefore did not ensure that these systems used it when needed.

Section 4.4, opening paragraph; Table 10 (288 dynamic queries per model) · Read the study

[c16] Table 4 reports 0.000 accuracy for zero-search-call cases on the dynamic split for all four models. These are conditional accuracies among no-call cases, not accuracies across all 288 queries; no uncertainty intervals or no-call subgroup counts are supplied in Table 4.

Table 4, Dynamic Split; columns 0 Calls, 1 Call, 2 Calls · Read the study

[c17] The manual failure analysis excluded missed retrieval: it examined only incorrect answers where search was invoked. The reported annotation samples were 150 incorrect static examples and 75 incorrect dynamic examples per model, with two annotators and Cohen’s kappa of 0.79.

Section 4.5, opening paragraph · Read the study

[c18] The manual retrieval-failure taxonomy includes only queries where web search was invoked and the final answer was incorrect. It therefore excludes failures to invoke search, and should not be presented as covering that distinct failure mode.

Appendix C, Failure Taxonomy Annotation, opening paragraph · Read the study