The practical brief
The finding
In one controlled AI-search benchmark, a two-stage defense cut the average rate of answers citing the manipulated document from 50.32% to 6.20% across five language models. [c1] [c2] [c3]
Why it matters
A citation uplift is a result inside a particular test. The business question is whether the evidence offered for a paid optimization service resembles the situation you care about.
| Experimental condition | Answers citing the attack document |
|---|---|
| No defense | 50.32% [c3] |
| GEO Defender | 6.20% [c3] |
What to try
BLURSOR’s practical interpretation
Ask a supplier to show the original and rewritten pages, the questions used, the comparison condition and repeated results. Treat this as our suggested due-diligence checklist; the study did not test a purchasing or marketing strategy.
Study boundary
The researchers tested an experimental defense against malicious rewrites. They did not test your website, ordinary content improvements, live AI-search products, referrals or sales.
One rewritten page, ten possible sources
The researchers started with groups of candidate documents and rewrote one page in each group to make it more likely to be cited. These malicious rewrites could preserve the original facts, so checking whether a page contained false statements was not enough to identify the manipulation.
Their defense worked in two places: it changed how candidate documents were ranked, then supplied guidance about using the selected sources in an answer. The resulting drop in attack-source citations is evidence about that experimental setup.
Ask for the test behind the sales pitch
Suppose a supplier shows you more citations after rewriting a page. Before deciding what that gain is worth, ask which questions were used, what changed besides the writing, and how the result compares with the original page under the same conditions.
Our interpretation is that those details belong in an evidence-based buying decision. This paper illustrates the difference between defended and undefended experimental conditions. It does not show that today’s AI-search providers use this defense, or that particular formatting choices will be penalized.
- Ask for a comparison you can inspect, rather than a percentage without context.
- Keep citation measurements separate from claims about visits, enquiries or sales.
- Do not conclude that adding useful references or clearer writing is itself manipulation.
Fewer attack citations did not settle answer quality
The authors also assessed answer quality with a language-model judge. Four of the five target models had confidence intervals that included zero, so those comparisons did not establish a systematic quality difference. GPT-5.5 showed a small mean decline, with its interval below zero.
The benchmark therefore supports a narrow conclusion: this defense reduced citation manipulation in this test. It does not establish a universal recipe for better answers, durable visibility or business growth.
Evidence [c8]
The source and its limits
A controlled generative-search benchmark: 616 attack-injected test instances, five language models and one rewritten document in each ten-document candidate pool.
- This benchmark does not measure deployed search products, website referrals or sales.
- Non-adversarial documents were a comparison group; that label does not independently establish credibility.
- Answer-quality comparisons used a model judge. The results do not establish a general improvement in answer quality.
Evidence behind this briefing
[c1] The test set contains 616 instances spanning all seven attack methods.
Experimental Setup, Dataset · Read the study
[c2] GEO Defender combines document reranking with guidance for using sources during answer generation.
GEO Defender, Framework Overview · Read the study
[c3] Across five target models, average attack-source citation rate fell from 50.32% without defense to 6.20% with GEO Defender.
Main Results; Table 1 · Read the study
[c8] Answer quality was assessed by a model judge. Four target models had confidence intervals including zero; GPT-5.5 showed a small decline.
Appendix C.1, Answer-Quality Evaluation · Read the study
[c4] Each attack-injected instance contains nine benign documents and one attack-rewritten document.
Experimental Setup, Dataset · Read the study