The practical brief
The finding
Competitor-aware rewriting improved controlled citation visibility, including on two datasets without retraining, while source-faithfulness and attribution scores fell. [c2] [c4] [c13] [c15]
Why it matters
Winning space in AI answers and preserving your source material are different objectives.
| Dataset and scope | Baseline | Competitor-aware rewrite |
|---|---|---|
| Original geo-bench | AgenticGEO: 27.95 ± 0.30 PAWC; 1,000 test examples | 32.62 ± 0.25 PAWC; 1,000 test examples [c2] |
| Synthetic competitive geo-bench | AgenticGEO: 25.20 ± 0.13 PAWC; 5,000 examples across five adoption levels | 29.93 ± 0.09 PAWC; 5,000 examples across five adoption levels [c2] |
| E-Commerce: transfer without retraining | AgenticGEO: 34.96 ± 0.53 PAWC; E-Commerce transfer test | 37.78 ± 1.10 PAWC; E-Commerce transfer test [c13] |
| Researchy-GEO: transfer without retraining | AutoGEO: 30.98 ± 0.71 PAWC; Researchy-GEO transfer test | 33.48 ± 0.24 PAWC; Researchy-GEO transfer test [c13] |
What to try
BLURSOR’s practical interpretation
Our interpretation: evaluate visibility and source fidelity together before scaling rewrites.
Study boundary
These are generated-answer experiments, not evidence of production visibility, factual accuracy or customer outcomes.
The decision: visibility without sacrificing source fidelity
Before commissioning broad AI-search rewrites, decide which claims, quotations and attributions must remain intact. This study finds higher citation visibility alongside weaker fidelity to original documents—not a measured decline in downstream answer accuracy.
The researchers studied generative engine optimization: rewriting content to increase its presence in language-model answers. Their question was whether editing choices should depend on competing documents, rather than applying the same formula everywhere.
How the competitive experiment worked
The method searched combinations of 15 rewriting strategies, then trained a smaller model to select combinations from a query, target document and competing documents. Selected instructions became simultaneous objectives in one rewrite prompt, not sequential edits.
Training started with 7,998 five-document queries, expanded to 39,990 examples across five synthetic competitor-adoption levels. Evaluation used the original 1,000-example benchmark and a 5,000-example competitive version. These levels were experimental conditions, not observed market adoption.
The selector saw all competing documents, but no adoption labels or optimization scores. That access gave it an information advantage over baselines without equivalent context.
Visibility gains transferred, but fidelity weakened
The main measure was Position-Adjusted Word Count, or PAWC: cited word count weighted toward earlier answer positions. It measures citation visibility, not correctness. With gpt-oss-120b generating answers, the method scored 32.62 versus AgenticGEO’s 27.95 on the original benchmark, and 29.93 versus 25.20 on the synthetic competitive benchmark.
The selector also transferred without retraining from geo-bench to E-Commerce and Researchy-GEO. It outperformed AgenticGEO on E-Commerce and AutoGEO on Researchy-GEO; the comparison below gives the measured values. These controlled gpt-oss-120b results report means ± standard deviations from one rewrite and five independent answer-generation evaluations.
Visibility nevertheless declined by 11.4% across synthetic adoption rates from 0 to 0.8—the slowest decline among tested methods.
Separately, model judges scored rewritten content against original sources. Faithfulness fell from 8.29 to 4.90 and attribution accuracy from 6.15 to 4.18 on 1–10 scales, while key-point coverage increased. These scores concern rewritten documents, not downstream answers.
A realistic interpretation for your content program
Transfer across two additional datasets suggests the selector was not useful only on its training benchmark. It does not establish a universal writing recipe or a production-search advantage. Our interpretation is to test competitor context as a variable, not treat it as a guaranteed tactic.
- Pilot a limited set of important pages before scaling; this is our suggested evaluation approach, not a validated tactic.
- Review rewritten claims, quotations and attribution alongside visibility changes.
- Record the competing documents included in each test; your business may lack the experiment’s full-corpus access.
What the study cannot tell your business
The experiment did not comprehensively simulate retrieval and re-ranking—the stages that select and order sources before answer generation. Neither benchmark gains nor transfer results establish more qualified visitors or sales.
The authors suggest fidelity losses might reflect verifiable enrichment rather than hallucination. Automated scores cannot establish that explanation or factual correctness, and longer documents can inflate coverage and citation measures.
The source and evidence
Study limitations and disclosures
- Generated answers and synthetic competitor adoption do not validate production behavior. Retrieval and re-ranking were not comprehensively modeled; the 15-strategy menu is not exhaustive.
- Full competitive-corpus access advantages the selector over baselines without equivalent context. Businesses may not reproduce that access.
- Fidelity scores use model judges, lack reported uncertainty and assess rewritten sources, not downstream answers. Document expansion can inflate coverage and citation metrics.
- Training explanations were generated after optimization. An automated audit flagged privileged-scorecard leakage in 113 of 12,124 traces (0.93%); those traces remained in training. No scorecard is supplied at inference.
- The authors are affiliated with Capital One, AI Foundations. They disclose AI assistance with writing, literature review and coding, and state that they originated and verified the work. The supplied source identifies no separate funding or specific commercial product interest.
- The authors acknowledge possible misuse through fabricated citations or unsupported claims. Prompt prohibitions against invented quotations, sources and unverifiable numbers do not establish output correctness.
Evidence behind this briefing
[c1] The method searches combinations of 15 rewriting strategies with BOCS, then trains a selector using teacher-generated reasoning and SFT+DPO. Training uses 7,998 five-document queries, synthetically expanded to 39,990 examples across adoption rates 0, 0.2, 0.4, 0.6 and 0.8; evaluation expands 1,000 test examples to 5,000. Competitor rewrites use random combinations, Quotation Addition and AutoGEO in roughly equal proportions when at least three competitors are rewritten. At inference, the selector sees the full competitive corpus but not adoption labels or BOCS scores.
§§3.1, 4.2–4.3.3; Figure 1 · Read the study
[c2] On the original 1,000-example geo-bench test corpus, the selector improves citation visibility rather than establishing factual accuracy or customer outcomes. PAWC is 32.62±0.25 versus AgenticGEO's 27.95±0.30 with gpt-oss-120b, and 29.55±0.15 versus 26.80±0.08 with Llama-3.3-70B-Instruct. On the 5,000-example synthetic competitive benchmark, averaged over five adoption rates, corresponding scores are 29.93±0.09 versus 25.20±0.13 and 25.87±0.03 versus 23.61±0.05. Tables report means and standard deviations from one temperature-0 rewrite and five temperature-0.7 engine evaluations; all rewriting uses gpt-oss-120b.
§5; §6, Main Results; Tables 1–2 · Read the study
[c3] In the synthetic competitor-adoption experiment with gpt-oss-120b, the authors report an 11.4% PAWC decline for their method across adoption rates 0–0.8, the slowest decline among tested methods. They report no statistically significant decline between 0 and 0.2, but the available passage supplies neither an uncertainty interval nor the significance-test details for that comparison. This is a controlled visibility result, not evidence about observed market adoption.
§6, Performance Scales with Adoption Rate; Figure 2 caption · Read the study
[c4] Visibility gains coexist with lower source fidelity: on original geo-bench with gpt-oss-120b, the method scores 7.48 for Key Point Coverage, 4.90 for Faithfulness and 4.18 for Attribution Accuracy, versus no-rewrite scores of 6.60, 8.29 and 6.15. These are LLM-judge scores on a 1–10 scale, reported as means without uncertainty. Its rewritten-document length ratio is 1.29. The authors interpret fidelity losses as potentially reflecting verifiable enrichment rather than hallucination, but these measurements do not establish factual correctness; Appendix B explicitly treats unsupported or fabricated claims as grounds for penalties.
§6, PAWC Optimization Introduces a Faithfulness–Enrichment Trade-off; Table 3; Appendix B.2–B.3 · Read the study
[c5] The evaluation uses generated-response proxies, synthetic competitor adoption and a non-exhaustive strategy menu rather than production-system validation. It does not comprehensively simulate retrieval and re-ranking. The selector also has an information advantage over baselines because it receives the full competitive corpus; the authors argue that competing documents are observable in practice.
Limitations · Read the study
[c6] Teacher reasoning is post-hoc, not the process that discovered successful combinations. An audit judged by gpt-oss-120b found privileged-scorecard leakage in 113 of 12,124 geo-bench training traces (0.93%): 61 of 6,062 chosen traces and 52 of 6,062 rejected traces. The paper describes training with this residual leakage and recommends filtering in future work; the student receives no scorecard at inference.
§4.3.2; Appendix A, Teacher Reasoning Trace Audit; Table 7 · Read the study
[c7] When multiple rewriting strategies are selected, the system combines their instructions into one prompt as simultaneous objectives rather than applying sequential rewrites. This describes implementation, not evidence that any individual strategy improves visibility.
Appendix E.3, Multi-Strategy Composition · Read the study
[c8] At both training and inference, the selector receives the query, all source documents, the marked target, and the strategy menu. Adoption rates, BOCS scores, and competitive-context labels are withheld; competitor document content remains visible.
Appendix F, Selector Input Prompt · Read the study
[c9] Teacher reasoning traces are generated by gemma-4-31B-it with access to BOCS scores, although the traces must ground explanations in document content and avoid directly mentioning scores. Chosen and rejected traces receive asymmetric scorecard information; their explanations should not be treated as independent validation of why strategies work.
Appendix G, Teacher Reasoning Trace Generation Prompt · Read the study
[c10] For the 7,998-example geo-bench training split, the authors report an average of 23 BOCS iterations per data point and 1,103,724 inference calls to generate preference pairs: each iteration uses one rewrite and five generative-engine calls. This is training-pipeline cost, not a visibility or accuracy result; deployment is described as one 2B selector call plus one rewrite pass.
Appendix H, Training Wall-Clock Time; Table 10 accompanying prose · Read the study
[c11] Artifact disclosure: the authors state that geo-bench, Researchy-GEO, and E-Commerce are publicly available under MIT licenses, that their use aligns with the creators’ intended GEO research use, and that the Gemma 4 models are Apache 2.0 licensed. These are source-reported licensing assertions, not independently verified legal conclusions.
Appendix I, Artifact Licenses · Read the study
[c12] Two representative query case studies illustrate changes in selector recommendations across competitor adoption levels. Each row presents competitor rewrite strategies, reasoning excerpts, and the resulting selection; at adoption rate zero there are no competitor rewrites. These examples describe adaptive recommendations, not measured visibility gains, factual accuracy, or causal effects.
Appendix K, Case Study: Adaptive Reasoning Under Competitive Pressure; Tables 12–13 · Read the study
[c13] In the controlled zero-shot transfer test, Competitor-Aware GEO achieved PAWC of 37.78 ± 1.10 on E-Commerce versus AgenticGEO’s 34.96 ± 0.53, and 33.48 ± 0.24 on Researchy-GEO versus AutoGEO’s 30.98 ± 0.71. These are means ± standard deviations using gpt-oss-120b, with one rewrite and five independent evaluations; the selector was trained only on geo-bench and applied without retraining.
§6, Table 4; columns E-Comm. and Researchy-GEO · Read the study
[c14] PAWC is a citation-visibility metric based on position-weighted cited word count, not a measure of factual accuracy or business outcomes.
Appendix B.1, PAWC definition, equation (5) · Read the study
[c15] The transfer results should not be presented as evidence of production visibility: the study uses LLM-generated response proxies and primarily controlled generation rather than a comprehensive retrieval and re-ranking pipeline.
Limitations · Read the study