The practical brief
The finding
Structured source renderings increased target citation counts by 0.50 markers per answer, but did not conclusively increase whether the source received any citation. [c2] [c3] [c5] [c12]
Why it matters
More citation markers and appearances in more answers are different outcomes; an optimization report can improve one without establishing the other.
| Outcome | Measured effect |
|---|---|
| Target citation count | +0.50 target markers per answer across 89 answer-target families; 95% CI +0.20 to +0.84 markers per answer. Survived multiple-test adjustment. [c2] [c7] |
| Target receives any citation | +4.5 percentage points across 89 answer-target families; 95% CI −1.4 to +10.4 percentage points. Inconclusive. [c3] |
What to try
BLURSOR’s practical interpretation
Our interpretation: measure citation entry and frequency separately, and repeat comparable tests before declaring a content change successful.
Study boundary
The experiment changed extracted text in frozen search transcripts, not live webpages, and measured visible credit rather than traffic, credibility or sales.
Decide which visibility outcome matters
Before paying to restructure content for AI search, distinguish appearing in more answers from receiving more citations within an answer. This study supports a narrow claim about additional credit—not a reliable route into the citation set.
The researchers studied citation allocation: how an answer engine assigns visible credit when several retrieved documents support the same fact. That happens after retrieval, so it is not a test of getting your page found.
The experiment changed text after retrieval
GPT-5.4 regenerated answers in 452 trials using frozen, multi-turn search transcripts. Researchers selected 113 competing document pairs without seeing ranks or answer outcomes, then crossed structured versus prose renderings with pair order. Effects were averaged across 89 answer-target families—the independent units used for analysis.
Blinded human review confirmed that 103 pairs supported the same required fact. This validated competing pairs, not rewrite fidelity. Grok predominantly generated the rewrites and checked fidelity in a separate pass; that was not independent model-family validation. A 23-treatment spot audit was AI-assisted, not independent human annotation.
More credit, unresolved entry, qualified rank effects
Structured rendering increased target citation counts by 0.50 markers per answer, surviving adjustment for multiple tests. Distinct citing sentences increased equally. Total citations did not reliably increase, and the matched competitor did not reliably lose credit.
The primary outcome—receiving any citation—rose an inconclusive 4.5 percentage points. Its 95% confidence interval ranged from −1.4 to +10.4 points; modest entry gains remain unresolved.
The observed rank-1 versus rank-5 citation gap was 42.3 percentage points. Controlled reordering produced +7.9 points in the main replay, failing multiple-test correction, and 0.0 points in 56 held-out pairs. These different populations and interventions do not establish a general ranking advantage.
One narrower rank result did survive correction: under structured rendering, higher rank increased the target’s share of citations within its competing pair by 10.5 percentage points (Holm-adjusted p=.0018). This covered only 72 answer-target families with at least one within-pair citation in both rank cells. Because it conditions on realized citations, it is descriptive—not evidence of an unconditional citation-entry or discoverability advantage.
Use measurement guidance, not a formatting recipe
Our interpretation is to evaluate presentation changes cautiously. Both renderings compressed and rewrote source text, so the experiment did not isolate formatting alone. A word-preserving list-marker test was unstable under repeat generation.
Repeatability matters: 15% of binary citation decisions changed across 120 regenerated cells. The aggregate count effect repeated, but on the same frozen cases—not an independent dataset.
- Report citation entry separately from citation count.
- Repeat comparable queries and show uncertainty rather than choosing a winner from one answer.
- Measure traffic or commercial outcomes separately; citations do not establish those benefits.
The source and evidence
Study limitations and disclosures
- Frozen transcripts, rewritten extracted text and a strict citation contract limit transfer to live webpages and sparsely citing systems. Results do not establish traffic, credibility or sales gains.
- Rank comparisons are not a causal decomposition. The positive rank-share result selects families with realized citations and cannot establish an unconditional entry advantage.
- Treatment validation remained model-dependent; human confirmation concerned competing pairs, not rewrite fidelity. Cross-model replication required citation repair for 121 answers and had missing cells.
- Archived inputs are search-returned snapshots, not complete webpages. Public code, prompts, hashes and derived analyses exclude licensed source text and raw provider traces, preventing fully self-contained source reproduction.
- The paper discloses AI assistance in query generation, pipeline implementation and sentence-level manuscript editing. The author reviewed the edits and accepts responsibility for the final text.
Evidence behind this briefing
[c1] The main experiment used GPT-5.4 to regenerate answers in 452 hash-verified trials crossing target rendering and within-call pair order: 113 machine-screened pairs across 89 independent answer-target families. Selection was outcome-blind; later blinded human adjudication confirmed 103 pairs as genuine same-proposition competitions. The estimand is conditional on the frozen transcripts and configuration, not other queries or retrievals.
Section 1; protocol and estimand details in Sections 4–5 and Appendices A–C · Read the study
[c2] Structured versus prose rendering causally increased family-weighted target citation count by 0.50 markers per answer (95% CI +0.20 to +0.84; Holm-adjusted p=.033), a secondary result surviving correction across 16 tests. Raw means were 2.69 versus 3.12 markers, not the family-weighted effect. Distinct citing sentences increased equally; total answer citations did not reliably increase and the matched competitor's change was −0.02 markers (95% CI −0.28 to +0.24). This is credit concentration, not demonstrated accuracy or reliance improvement.
Section 6.1; Table 2 · Read the study
[c3] The prespecified primary outcome—whether the target received any citation—was inconclusive: +4.5 percentage points (95% CI −1.4 to +10.4; p=.168) across 89 families. The design had roughly 80% power only for effects around 8.5 points or larger. Restricting to 103 human-confirmed pairs yielded +5.9 points (p=.066), still not confirmed source admission.
Section 6.1; Table 2; Limitations—Power and multiplicity · Read the study
[c4] The descriptive first-exposure rank-1 versus rank-5 citation gap was 42.3 percentage points (85.1% of 276 documents versus 42.8% of 311). Controlled pair reordering yielded +7.9 points in the main replay (95% CI +1.1 to +14.9; Holm p=.350) and 0.0 points in 56 held-out pairs (95% CI −5.4 to +5.4). These differing populations, text regimes and channels do not support a causal decomposition or general ranking advantage.
Section 6.3; Appendix A first-exposure rates; Table 2 · Read the study
[c5] Fresh decoding of the same frozen cases exposed substantial measurement noise: binary target-citation decisions agreed in 85% of 120 regenerated cells, exact target counts in 61.7%, and exact family effects in 17 of 30 families. A method-of-moments estimate assigned 45% of single-draw family-effect variance to decoding. Aggregate count effects repeated at about +0.55 markers, but this was not independent-dataset replication.
Section 6.5; aggregate repeatability qualification in Section 6.1 · Read the study
[c6] The treatment was a jointly generated rewrite package, not formatting alone or an edit to the original live webpage. Both arms compressed the source and differed lexically. The word-preserving list-row ablation reversed direction in the repeated subset, precluding a stable list-marker recipe; the strict citation contract and visible-credit outcome further limit marketing interpretations.
Limitations—Construct and treatment; Section 6.4 · Read the study
[c7] The experiment observes all four text-by-rank cells per pair and estimates within-pair contrasts, averaging effects into 89 answer-target families from 113 pairs. Family-bootstrap intervals and sign-flip tests support inference; the primary replay uses GPT-5.4 with one independently sampled generation per cell.
Appendix G, Estimators and Inference; model configuration in Appendix L · Read the study
[c8] Among 16 declared secondary tests in the scaled wave, structure increased target citation count by 0.50 markers per answer (95% CI +0.20 to +0.84; Holm p=.033). Higher rank increased within-pair citation share under structured rendering by 10.5 percentage points (Holm p=.0018), but that share analysis conditions on 72 families with citations in both rank cells and is descriptive, not an unconditional discoverability estimate.
Appendix H, Table 8 and Count and share baselines · Read the study
[c9] The mechanical ablation preserves word sequence and changes only line breaks and list markers across 113 targets. Its full-cohort estimate is +6.7 percentage points, but repeat generations in 30 hash-selected families yield negative, uncertain estimates. This limits a firm mechanism claim that formatting alone reliably explains the effect.
Appendix I, Ablation repeatability; manipulation specified in Appendix F.3 · Read the study
[c10] Rank effects are not stable across tested regimes: the 56 held-out pairs show near-zero displacement-stratum effects, including full-window swaps. The authors identify original versus rewritten target text and sampling width as operative differences; scaled-wave rank subgroups are exploratory and uncorrected.
Appendix J, Held-out confirmation geometry; Table 10 · Read the study
[c11] Grok-4.3 replication covers 219 of 224 planned cells and 54 complete pairs out of 56. Worst-case completion bounds the repaired-protocol structure effect at +2.7 to +7.1 percentage points. This is a repair-dependent replication: 121 answers needed citation repair, and the larger scaled Grok replay was aborted.
Appendix K, Cross-Model Replication and Missingness; repair count in Appendix D.8 · Read the study
[c12] The intervention changes serialized search-result text after retrieval and extraction, not live publisher pages or visual design. Exa retains textual organization but discards typography, visual rendering, and DOM structure; whether publisher-side edits survive that pipeline is outside the design.
Appendix L, How the retriever serializes document structure · Read the study
[c13] Grok-4.3 generated the treatments and checked fidelity in a separate blinded pass; this was procedural separation, not independent model-family validation. The 23-item treatment audit was AI-assisted, not independent human annotation. Human confirmation concerned shared-evidence competing pairs (103/113), not treatment fidelity, which retains model-dependent error.
Limitations—LLMs in the measurement loop · Read the study
[c14] Appendix C specifies a GPT-5.4 fallback for treatment-generation protocol failures: one accepted treatment used that fallback, and excluding it left the primary estimate unchanged (+4.55 percentage points over 88 families). Thus treatment generation was predominantly, not exclusively, Grok-4.3.
Appendix C—Treatment Generation and Integrity · Read the study
[c15] The archive contains search-returned snapshots, not complete webpages. Licensing restrictions on third-party text and raw provider traces prevent fully self-contained source reproduction. The public artifact provides code, prompts, the pipeline, hashes, derived records, and no-network analysis regeneration, but these do not eliminate the independent source-audit limitation.
Limitations—Data and reproducibility · Read the study
[c16] The author discloses AI assistance in query-set generation, pipeline implementation, and sentence-level manuscript editing after the first draft; the author reviewed the edits and accepts responsibility for the final text.
Ethical Considerations—AI assistance disclosure · Read the study
[c17] The configuration identifies Azure Grok-4.3 as the model used for evidence review, treatment audit, cross-model replication, and citation alignment. This excerpt establishes model involvement in treatment auditing, but does not establish who generated treatments or the independence of fidelity checks.
Appendix L — Configuration and Model-Visible Interface, Models and decoding · Read the study
[c18] The reported human confirmation concerned whether competing documents genuinely supported shared evidence. In the walkthrough, the adjudicator treated the shared cost unit as decisive; this pair belonged to 103 confirmed competitions, while ten failed pairs were excluded in sensitivity analysis. This is pair-level adjudication, not evidence of independent human treatment-fidelity validation.
Appendix E — End-to-End Walkthrough: One Query, Stage 4: blinded evidence audit · Read the study
[c19] The public artifact provides pipeline and analysis code, prompts, configurations, selection rules, derived records, numerical summaries, and frozen-input hashes, with no-network regeneration of the consolidated analyses. Licensed third-party page text and raw provider traces are excluded, so the artifact does not provide fully self-contained reproduction of source inputs; hash verification requires licensed re-acquisition.
Appendix M — Archives, Artifact, and Amendments, Artifact · Read the study