The practical brief
The finding
Across three open-weight models and 247 English queries, a purpose-built detector reduced average weighted attack success from 55.7% to 29.2%; general safety filters offered much smaller reductions. [c1] [c2] [c3] [c4] [c6]
Why it matters
A business using AI to summarize retrieved sources should not equate a safety filter with protection against false factual claims.
| Defense | Three-model average weighted attack success |
|---|---|
| No defense | 55.7% weighted score; 247 queries per model across three models [c1] [c2] |
| Granite Guardian | 54.0% weighted score; 247 queries per model across three models [c3] |
| Llama Guard 3 | 52.5% weighted score; 247 queries per model across three models [c3] |
| Purpose-built detector | 29.2% weighted score; 247 queries per model across three models [c4] |
What to try
BLURSOR’s practical interpretation
Our interpretation: evaluate factual distortion and non-attacked answer performance together before relying on a defense.
Study boundary
The study did not test commercial AI-search products, publisher visibility, customer belief or business outcomes.
The decision: what does your AI safeguard actually check?
If your business uses AI to turn retrieved documents into answers, the risk is not limited to offensive content or explicit malicious instructions. An ordinary-looking source can contain a false claim that the model repeats. This study tests that risk—not whether optimizing your website earns more AI citations.
Generative engine optimization means editing content for use in AI-generated answers. The researchers distinguish fact-preserving edits from edits that insert targeted misinformation. Their question is whether defenses can block the latter without unnecessarily suppressing the former.
The study changed one source, not the whole web
The benchmark contains 247 English queries with human-checked attack criteria and ground-truth materials. For each query, researchers changed one of five source documents while leaving four unchanged. Rewrites passed a quality screen intended to keep plausible, natural-looking content.
A controlled retrieval-and-answer pipeline tested Gemma, Qwen and Llama open-weight models. An AI judge scored attack success as 1 for an unhedged false assertion, 0.5 for partial success, and 0 for failure. The reported attack success rate is therefore a weighted average, not the percentage of fully successful attacks.
Specialized filtering helped; general guardrails mostly did not
Without a defense, average weighted attack success was 55.7%. Granite Guardian’s main-text reduction of 1.7 percentage points was not statistically significant; Llama Guard 3 reduced it by 3.2 points.
NeMo’s apparently strong Llama result came from refusing almost everything, including 98.4% of clean-condition answers. That model was excluded from NeMo’s main-table average, so its average is not directly comparable with the three-model results.
The authors’ detector, C-GEO Guard, reduced weighted attack success to 29.2%: a 26.5-percentage-point absolute reduction, or 47.6% relative. Relevance, completeness and clarity stayed nearly unchanged. Those quality scores do not establish factual correctness, and substantial residual distortion remained.
A realistic interpretation: test protection and lost utility together
Our interpretation is that businesses should ask what a safeguard detects, rather than treating “guardrails” as a single capability. Policy-violation screening and factual-misinformation detection are different tasks.
Transfer tests reinforce the trade-off. On GPT 5.5 rewrites with Qwen answering, the detector reduced attacks but lowered fact-preserving-condition accuracy by 4.2 percentage points. An independently written attack template produced a smaller relative reduction. Neither result proves reliable deployment across arbitrary sources.
- For an internal pilot, compare attacked and non-attacked inputs; a lower attack score achieved through blanket refusal is not useful protection.
- Track correctness separately from clarity, and inspect wrongly blocked source passages.
- Treat these as evaluation priorities, not proven tactics for gaining public AI visibility.
What the study cannot tell your business
No proprietary APIs or end-to-end commercial search products were tested. English-only queries, single-document attacks and topic selection constrain generalization. Coordinated attacks, attacks outside the defined categories and attempts designed to evade the detector remain untested.
The findings support a narrower conclusion: specialized defenses deserve evaluation in controlled source-summarization workflows. They do not establish protection for a particular search platform, improved recommendations for your brand or changes in audience trust.
The source and evidence
Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization
Study limitations and disclosures
- The evaluation covers 247 English queries and three open-weight models, not commercial AI-search products. Small defense differences may be underpowered; multilingual, coordinated multi-document, open-set and detector-aware attacks were not evaluated.
- Quality screening altered the representation of topics, limiting application of aggregate results to any particular business sector.
- Attack scoring used an AI judge. Human validation covered 150 undefended model–query instances spanning 116 unique queries, not every metric or defense condition.
- Transfer did not consistently preserve non-attacked performance: the GPT 5.5 rewrite test lowered fact-preserving-condition accuracy by 4.2 percentage points, and an independent template produced weaker mitigation.
- Experiments used all 247 instances, but the public release withholds 10 sensitive misinformation rewrites and excludes original source documents. Data and detector weights use gated CC BY-NC 4.0 licensing; code uses Apache 2.0. Commercial reuse requires checking these terms.
- Bing Zheng and Wenming Yang are affiliated with Tsinghua University’s Shenzhen International Graduate School; Zongyao Zhao is affiliated with the University of Hong Kong’s Department of Electrical and Computer Engineering. The paper reports partial funding from Shenzhen strategic-emerging-industries foundations and the Shenzhen-Tsinghua AI fundamental and frontier research project. It does not report commercial conflicts of interest. The authors disclose GPT language polishing, Claude Opus coding assistance and Gemini assistance with rebuttal clarity, and state that they developed and verified the research claims and conclusions.
Evidence behind this briefing
[c1] The paired benchmark distinguishes legitimate, fact-preserving GEO from misinformation: one of five source documents changes while the other four remain unchanged. ASR is an LLM-judged weighted score—full success 1.0, partial success 0.5 and failure 0.0—not simply the percentage of fully successful attacks.
§3.4 Evaluation Metrics, Attack Success Rate; paired construction and source controls in §§3.2–3.3 · Read the study
[c2] In the controlled 247-query harness, undefended average weighted ASR across the three victim models was 55.7%, with a 95% bootstrap CI of 53.1–58.2%. This measures answer distortion from quality-gated single-document attacks, not real-world discoverability or audience belief.
§5.2 Attack Mitigation Results, paragraph following Table 3 and Figure 3 · Read the study
[c3] Off-the-shelf guardrails offered little protection in this harness: Granite's reported 1.7-percentage-point reduction was not statistically significant; Llama Guard 3 reduced ASR by 3.2 points. NeMo's apparent success on Llama-4 was a refusal artifact, so that model was excluded from its main-table average.
§5.2 Attack Mitigation Results; Table 3 exclusion footnotes; Appendices G and J · Read the study
[c4] The study's 184M-parameter GEO-specific detector reduced three-model average ASR from 55.7% to 29.2%: 47.6% relative, or 26.5 percentage points absolute. The paired reduction's 95% CI was 23.8–29.3 points. ID answer quality stayed nearly unchanged on the separate relevance/completeness/clarity scale; this is not evidence of perfect factual accuracy.
§5.2 Attack Mitigation Results, C-GEO Guard paragraph; Table 3 and Appendix G · Read the study
[c5] Transfer to GPT 5.5 rewrites retained attack mitigation but worsened benign utility: with Qwen as victim on all 247 rewrites without quality filtering, ASR fell from 55.7% to 22.1%, while legitimate IP accuracy fell 4.2 percentage points and chunk-level IP false positives reached 2.80%. A separate independently written template produced a smaller 32.0% relative ASR reduction, so transfer strength depends on the tested setting.
§5.4 Cross-Rewriter Transfer; Appendix F, Table 13; independent-template qualification in §6.1, Table 7 · Read the study
[c6] The study does not establish protection for commercial AI-search products. It tests only three open-weight models, 247 English queries and single-document control; smaller defense differences may be underpowered. Multilingual, coordinated multi-document, open-set and detector-aware attacks remain untested.
Limitations, Victim model coverage; Scale and scope; Adaptive and open-set attacks · Read the study
[c7] Quality-gate selection changes topic representation, limiting how directly aggregate results apply to a particular business sector. The 247-query set covers 17 categories, but naturalness and refusal patterns over-represent some topics and under-represent others.
Appendix C Query Topic Distribution, introductory paragraph and Table 9 · Read the study
[c8] The experimental data and public release differ: experiments used all 247 instances, but 10 sensitive misinformation rewrites are withheld publicly and original source documents are excluded. Benchmark data and guard weights are gated under CC BY-NC 4.0, whereas code uses Apache 2.0; commercial reuse requires attention to those terms.
Ethics Statement; repository URLs in Footnotes 1–2 · Read the study