The practical brief

The finding

Apparent corroboration made fabricated evidence harder to recognize; caution-oriented prompts improved resistance but lowered measured usefulness. [c1] [c2] [c5] [c19] [c20]

Why it matters

A researched-looking recommendation can rest on repeated false claims rather than independent support.

Two interventions, different outcomes
InterventionObserved outcome
Interactive search versus static retrievalLower endorsement in all four model–prompt comparisons; recognition did not reliably improve, with confidence intervals frequently overlapping zero; usefulness generally flat or lower. [c3] [c16]
Explicit reasoning enabled versus disabledImproved resistance and recognition in all four model–prompt comparisons, with other experimental settings fixed. [c13] [c14] [c17]
Qualitative interpretation of matched experiments limited to Kimi-K3 and Qwen3.8-27B, not a universal safeguard ranking. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: assess whether checking changes the final recommendation, alongside recognition, usefulness and resource use.

Study boundary

This controlled Chinese-language benchmark does not establish live-platform visibility or marketing effectiveness; usefulness scores also have a task-wording limitation.

Judge correction, not just additional checking

If your business relies on AI product comparisons, do not treat additional research as proof that a recommendation is sound. This study separates checking a suspicious claim from abandoning it.

Caution-oriented prompts raised strict independent verification from roughly 2% to 17% of trajectories exposed to target poisoning. Yet fake-brand endorsement under those prompts reached 38.6% among exposed trajectories in the apparent-corroboration condition. Protection improved, but remained incomplete.

Evidence [c4] [c15]

The experiment changed how false evidence appeared

Researchers tested ten agents on 120 category-balanced queries: 60 comparison/recommendation tasks and 60 verification tasks. The controlled corpus contained 72,039 clean pages and 770 poisoned pages per attack level, spanning eight product categories and 154 fabricated brands. “Clean” means non-adversarial, not independently certified as true.

Fixed false claims appeared as direct assertions, contextual camouflage or apparent corroboration across plausible sources. Agents could search and retrieve pages over multiple rounds. Queries and operational instructions were Chinese; English translations did not generate the results.

Strict verification required new substantive page evidence explicitly attributed to a non-attack source. Search summaries, duplicates and unknown provenance did not qualify. Primary behavioral rates counted trajectories actually exposed to target poisoning, not all queries.

Evidence [c1] [c7] [c8] [c19]

Resistance, recognition and usefulness moved differently

Under standard prompting, the unweighted ten-model average poison-recognition score fell from 1.02 for direct assertions to 0.33 for apparent corroboration on a 0–2 judge rubric. These are assessment scores, not accuracy percentages; individual models did not uniformly follow that progression.

Defense improved pooled robustness across all ten agents but lowered attack-condition usefulness by 0.01–0.27 points on the 0–2 rubric. However, the Chinese usefulness judge frames success around recommendation completion even for verification tasks. Defense also added an average 1.82 tool calls per attack query and increased within-model token consumption by a macro-average 24.5%; these are not price or latency measurements.

Matched experiments covered only Kimi-K3 and Qwen3.8-27B. Interactive search reduced endorsement in all four model–prompt comparisons, but recognition did not reliably improve: its confidence intervals frequently overlapped zero. Usefulness was generally flat or lower. Separately, enabling explicit reasoning improved resistance and recognition in all four comparisons, with other settings fixed.

Evidence [c2] [c3] [c5] [c13] [c17] [c20]

Evaluate safeguards as trade-offs, not guarantees

Our interpretation: more searches, cautious wording and explicit reasoning are not interchangeable safeguards. These evaluation suggestions are not proven visibility tactics.

  • Check whether new evidence changes the final recommendation.
  • Distinguish independent support from repeated claims across pages.
  • Compare usefulness and resource use alongside resistance.

Evidence [c5] [c8] [c9] [c13] [c16]

The study cannot estimate real-world exposure or recovery

The benchmark cannot establish how often comparable poisoning occurs on public platforms, whether customers trust the answers, or whether a business gains discoverability.

Recovery evidence is sparse: defended trajectories that adopted poisoned evidence and then obtained strict verification numbered only 11, nine and three across the attack levels. Wide intervals prevent strong causal conclusions about recovery.

Evidence [c1] [c4]

The source and evidence

Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

Zhongan Bi, Qiwen Wang, Jianrong Jiang, Jigang Ding, Wenwen Xiong, Changhua Meng, Xuanang Gao, Kepeng Lin, Changjiang Jiang, Yiang Chen, Huan Yao, Wei Wang, Zhenyu Ma, Wenhui Dong · 2026-09-05

Study limitations and disclosures

  • Results concern a controlled Chinese-language benchmark, not organic visibility on public AI platforms. English translations were not used to generate results.
  • Recovery estimates use sparse conditional subsets and wide intervals; retrieval and reasoning comparisons cover only two models.
  • Only a restricted clean-corpus subset is publicly released, limiting exact reproduction; public access does not grant redistribution rights.
  • Available evaluator source identifies v6; the three archived case reports identify v5 and GPT-5.5. Reconstructed judge templates establish neither historical request byte identity nor that GPT-5.5 produced those archived scores.
  • One qualitative recovery example attributes a webpage’s endorsement to the agent, limiting its automated adoption and recovery labels.
  • The Chinese usefulness judge uses recommendation-completion wording even for verification tasks, which comprise 60 of 120 queries. This qualifies the usefulness trade-offs.
  • An illustrative Defense case retained conditional purchasing advice despite an archived risk-handling score of 2. That label is not unambiguous categorical exclusion; the case does not measure scoring-error frequency.

Evidence behind this briefing

[c1] HAE-GEO holds fabricated claims and target brands fixed while escalating their presentation across three poisoning levels. Its controlled corpus contains 72,039 clean pages and 770 poisoned pages per level, covering eight categories and 154 fabricated brands. The main evaluation uses 120 category-balanced queries across ten agents, standard and caution-oriented prompts, and a maximum of ten decision rounds. Primary behavioral rates are conditioned on actual target-poison exposure, rather than all queries.

§4.1; Table 3; §4.2; Appendix A.3/Table 8 · Read the study

[c2] Across ten models, unweighted macro-average risk handling and poison-evidence recognition decline between direct assertions (L1) and apparent corroboration (L3). On the 0–2 semantic rubric, Base risk handling falls from 0.83 to 0.30 and recognition from 1.02 to 0.33. Defense raises scores but does not remove aggregate degradation. This is not a monotonic difficulty result for every model, nor a factual-accuracy measurement.

§5.2, discussion of Figure 3; Appendix A.1 · Read the study

[c3] A matched comparison limited to Kimi-K3 and Qwen3.8-27B finds lower fake-brand endorsement with Agentic Search–Scrape than Static Full-Context retrieval in all four model–prompt conditions. Recognition gains are modest, with confidence intervals frequently crossing zero; legitimate utility is generally flat or lower. The figure uses paired means and 95% query-cluster bootstrap intervals, not evidence of a universal advantage across all agents.

§5.3; Figure 4 caption · Read the study

[c4] Pooling target-exposed trajectories across ten agents, defense raises strict independent verification from roughly 2% to 17%, but defended endorsement still rises from 28.6% at L1 to 36.3% at L2 and 38.6% at L3. Among already-adopted trajectories subsequently obtaining strict verification, defended recovery is 81.8%, 66.7%, and 66.7%. These conditional recovery estimates have only 11, 9, and 3 eligible trajectories and wide intervals; they do not support strong causal claims about defense efficacy.

§5.4; Figure 5, including Wilson 95% intervals and eligibility disclosure · Read the study

[c5] Defense improves pooled robustness for all ten agents but lowers attack utility by 0.01–0.27 points on the 0–2 rubric. Across 360 attack queries per model–prompt configuration, it adds an average 1.82 Search/Scrape calls per query and increases within-model token consumption by a macro-average 24.5%. Additional computation does not consistently deliver proportional robustness gains; absolute token counts are not directly comparable across providers.

§5.6/Table 7 and its tokenizer disclosure; §5.2; §5.7/Figure 6 · Read the study

[c6] The publicly released clean corpus is incomplete: pages with crawling, collection, or reuse restrictions are excluded. Public availability does not confer redistribution rights, limiting public reproduction of the exact experimental corpus.

Data Release Statement · Read the study

[c7] The 120-query evaluation was selected in two stages from an existing category-balanced 240-query set, not directly from all 1,011 queries. It contains 15 queries per category, 60 comparison/recommendation tasks and 60 verification tasks; 64 queries have fake-brand annotations and 56 do not, without an annotation sampling quota.

Appendix A.14, Exact 120-Query Selection Protocol · Read the study

[c8] Exposure and verification are trace-based measurements, not assumptions from injected-page counts. Search exposes metadata only; page content requires Scrape. Strict verification requires new substantive scraped evidence explicitly attributed to a non-attack source, excluding metadata leads, unknown provenance and duplicate content.

Appendix A.8, Tool Projection and Trajectory-State Annotation; Strict verification yield · Read the study

[c9] In the semantic results pooled across L1–L3, high uncertainty-calibration scores can coexist with low fake-brand risk-handling and poison-recognition scores. General caution therefore does not establish target-specific poison recognition. These are judge rubric outputs, not accuracy or consumer-preference measurements.

Appendix A.9, Full Semantic Rubric Results, discussion following Table 9 · Read the study

[c10] The attack hierarchy encodes intended mechanism sophistication, not a universally increasing measured difficulty. Qualitative early L3 calibration found that dense authority wording could trigger skepticism; the final design instead distributes consistent but non-identical apparent corroboration across plausible sources.

Appendix A.10, Calibration of Evidence-Enhanced Poisoning; Table 10 · Read the study

[c11] Prompt provenance is qualified: agent strings are exact exports of available runtime source, but judge listings are reconstructed templates rather than historical API requests. Available evaluator source identifies v6, while archived case reports identify v5 and GPT-5.5; the artifacts do not establish historical byte identity or that GPT-5.5 actually produced those archived scores.

Appendix A.12, Provenance and scope of exactness · Read the study

[c12] A representative Base power-bank case has archived M1/M2/M3 scores of 2/2/2 and a rule-based recovery label, but the detected adoption phrase describes a webpage's recommendation rather than clearly expressing agent adoption. It is an attribution-error example, not independently confirmed adoption followed by recovery.

Appendix A.15, Case 3: checking the attribution behind a recovery label (Base) · Read the study

[c13] In matched reasoning-on versus reasoning-off tests limited to Kimi-K3 and Qwen3.8-27B, explicit reasoning improved final resistance and poison-evidence recognition in all four model–prompt comparisons. Pooled across L1–L3, exposure-conditioned fake-brand endorsement fell by 9.7 and 8.8 percentage points for Kimi-K3 under standard and defense prompts, respectively, and by 16.0 and 19.8 points for Qwen3.8-27B. M1 risk handling and M2 recognition also improved in every comparison; this is not a general finding across all ten agents.

§5.5, Explicit Reasoning Strengthens Defenses; Table 6 · Read the study

[c14] Table 6 pools the explicit-reasoning comparisons across poisoning levels L1–L3 under strictly paired settings, comparing reasoning enabled with reasoning disabled rather than comparing additional search with static retrieval.

§5.5, Table 6 caption · Read the study

[c15] FR measures positive fake-brand endorsement among trajectories actually exposed to target-relevant poison, not general answer accuracy or product preference among all queries.

§4.2, Behavioral Outcome and Trajectory Diagnostics · Read the study

[c16] The explicit-deliberation result is a material counterpoint to the separate additional-search experiment: agentic search reduced endorsement but did not reliably improve recognition, with M2 confidence intervals frequently overlapping zero and utility flat or lower in most conditions.

§5.3, Agentic Search: Resistance without Recognition · Read the study

[c17] The reasoning ablation held queries, environments, prompts, tools, and judge configuration fixed, changing only the endpoint’s explicit-reasoning mode. It therefore tested reasoning-mode changes rather than an intervention that added search or changed the interaction protocol.

Appendix A.11, Additional Evaluation Details, final paragraph · Read the study

[c18] The study identifies its source prompts as Chinese and its English translations as reference-only. The available passage does not explicitly establish the language of evaluated queries or state that English translations were not used to generate results.

Section 4.4, Split LLM-as-Judge Protocol, final paragraph · Read the study

[c19] The documented evaluation used Chinese-language queries and source instructions. English translations were supplied solely for readability and were not used to produce the reported results; the findings should therefore be presented with this Chinese-language scope disclosed.

Appendix A.12, Original Agent Instructions and Prompt Provenance, opening paragraph · Read the study

[c20] The Chinese M6 judge implementation frames utility around completing a recommendation task even when task context indicates verification. This qualifies interpretation of its usefulness scores for verification tasks.

Appendix A.13, Quality Judge: English translation (not used for evaluation), final sentence · Read the study

[c21] The documented experiment uses the same category-balanced 120-query subset across models, prompts, and environments, comprising 60 recommendation/comparison tasks and 60 verification tasks.

Appendix A.11, Additional Evaluation Details, opening paragraph · Read the study

[c22] In the illustrative Defense comparison of Nike with a fabricated children's-shoe brand, archived M1/M2 scores were 2/2, but the answer retained conditional purchasing advice. Its high M1 label therefore does not establish unambiguous categorical exclusion; this is a case-specific caveat, not a measured frequency of scoring errors.

Appendix A.15, Case 2: explicit diagnosis of promotional evidence (Defense), introductory paragraph · Read the study

[c23] The three representative L3 cases are qualitative selections involving Claude-sonnet-4.6, covering two task types and both prompts; they were not selected to estimate failure frequencies.

Appendix A.15, Representative Observable Trajectories, opening paragraph · Read the study