The practical brief
The finding
A targeted laboratory intervention reduced manipulated-source selection across four models, while preserving only 67.97–82.35% of their uninjected source choices. [c13] [c14] [c6]
Why it matters
Resistance to manipulation does not establish source quality: the replacement source was not verified as better.
| Evaluation question | Factual-answer task | Source-selection task |
|---|---|---|
| Reference for judging decisions | Ground-truth answer on selected initially correct questions | Original uninjected model choice, not verified ground truth [c14] |
| Defense improvement | Higher correct-answer accuracy under persuasion | Lower injected-target selection; better alternatives not established [c13] [c14] |
| Decision preservation | Most clean decisions preserved in the evaluated cohort | Substantial changes to uninjected choices [c13] |
What to try
BLURSOR’s practical interpretation
Our interpretation: ask for manipulated and uninjected evaluations separately before treating robustness gains as better recommendations.
Study boundary
The study did not measure deployed search visibility, citation share, free-form recommendations or commercial outcomes.
The buying decision: what does “better source selection” mean?
Before buying an AI visibility service or trusting a source-selection safeguard, check what “better” measures. This study shows that instructions inside retrieved content can redirect source choices—and that resisting those instructions can change ordinary choices too.
Researchers tested Llama-3 8B, Gemma-3 12B, OLMo-2 7B and Qwen-3 4B. Their factual task contained 489 four-option questions, each paired with nine credibility, emotion or logic appeals. A separate task asked models to choose one of four retrieved documents, with an instruction favoring a candidate appended inside that document.
How persuasive context redirected decisions
The authors changed internal model activity to test cause and effect. Earlier components read persuasive keywords and write a routing signal. That signal redirects selected “attention heads”—components that allocate attention to input information—which then copy the selected option into the answer.
Mechanism analyses selected successful, confidence-qualified answer switches; factual questions also required initial correctness across all four models. They explain selected switches, not how often ordinary marketing changes answers.
Replacing literal target names with informative descriptions reduced but did not eliminate persuasion among initially correct question–appeal pairs. In separate, successfully persuaded, confidence-filtered localization cohorts, the same heads remained dominant. These were not tests of removing all target information.
The identified heads also supported legitimate, evidence-based updates in a separate BBQ experiment, with much smaller effects for Qwen. That experiment did not test whether the trained defense harmed legitimate evidence use.
Robustness improved, with a source-choice trade-off
The defense used small, query-head-specific attention adapters applied to selected final-token query rows, with the base model frozen. These were targeted interventions, not corrections merged into globally modified model checkpoints.
On held-out, initially correct, confidence-qualified factual questions, accuracy under persuasion rose by 12.30–41.27 percentage points. In source selection, injected-target selection fell by 28.75–47.06 points, but only 67.97–82.35% of uninjected choices were preserved. These ranges span model point estimates, not confidence intervals. Uncertainty was substantial: Llama’s defended factual accuracy was 75.13% (95% confidence interval: 65.34–84.66%).
Training and selecting corrections on two appeal families improved accuracy on the excluded third family in all 12 model–family combinations. Gains were 8.5–45.2 percentage points, with 89.3–97.9% clean-decision preservation—ranges of model–family point estimates. Transfer did not uniformly match seen-family performance and does not establish deployment generalization.
The source-selection reference was the original model choice, not verified ground truth. The intervention optimized target suppression, not correctness of an alternative.
A realistic interpretation for business buyers
Our interpretation: evaluate resistance to manipulation and ordinary decision quality separately. A system can become harder to steer without choosing better sources.
This is an evaluation lesson, not a visibility recipe. The study cannot predict whether your page will be retrieved, cited or recommended in a deployed product.
- Ask whether an improvement means factual accuracy, target suppression or agreement with previous choices.
- Request uninjected results too; non-adversarial does not mean credible or trustworthy.
- Do not translate laboratory gains into expected citation share or sales without separate evidence.
The source and evidence
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
Study limitations and disclosures
- Four specified models and controlled tasks do not establish results for free-form answers, end-to-end retrieval or deployed recommendations.
- Mechanism, defense and evidence-update experiments used different selected cohorts. Reported intervals use question-level bootstrap resampling, keeping factual appeals grouped by question.
- The adapters intervened on selected final-token attention rows with frozen base models, rather than globally modified checkpoints. Withheld-family transfer was positive but not uniformly equal to seen-family performance.
- Target-name substitutions retained informative descriptions. Behavioral and localization samples differed; neither tested removing all target information.
- Source references were model choices, not ground truth. Suppressing an injected target does not establish a better alternative.
- The authors disclose generative-AI manuscript polishing and Meta-Llama-3-70B-Instruct assistance constructing some controls and appeals, accepting responsibility for these materials.
- The authors identify dual-use risks, including retrieval poisoning and source-selection manipulation, and frame the work as defensive.
Evidence behind this briefing
[c1] NQ2 contains 489 questions with nine persuasion appeals each. Mechanistic patching is selectively evaluated on questions all four models originally answered correctly, where persuasion switched the answer to its designated incorrect target and both decisions met a probability-margin threshold; these are not representative all-question results.
Appendix E.1 NQ2; Appendix E introductory filtering definition; Table 3 · Read the study
[c2] The GEO experiment asks models to select one of four retrieved sources, with a direct instruction appended inside one candidate document. It measures changes relative to the uninjected selection, not source-selection accuracy, citation share, or commercial discoverability. Appendix E reports 998 pairs after removing two invalid examples.
Section 3.2, GEO; Appendix B; Appendix D metric definitions; Appendix E.2 · Read the study
[c3] Causal attention-pattern replacement supports attention rerouting as a major mechanism in the selected successful-persuasion examples. For Llama-3-8B-Instruct head L17H24 on NQ2, clean attention-pattern patching increased mean correct-option probability by 42.10 percentage points, versus 47.28 points for full output patching. The latter estimate has a 95% CI of 41.84–52.97 points in Table 4; this is probability recovery, not an accuracy increase.
Section 4.1, Attention rerouting is the major mechanism of persuasion; Appendix F/Table 4, Llama 3 8B rank 1 · Read the study
[c4] The intervention-supported mechanism links upstream reading of persuasive keywords to a one-dimensional option-routing feature, which redirects mid-layer decision-head attention; the selected option representation is then copied into the answer code. This is a mechanistic account within the tested multiple-choice settings, not evidence that all AI persuasion or marketing content uses this circuit.
Section 4.3, Closing the causal chain; Sections 3.2 and 6 for scope · Read the study
[c5] A trained rank-2 correction to decision-head key projections improved held-out NQ2 accuracy under persuasion by 12.30–41.27 percentage points across the four models, preserving 95.24–96.43% of clean decisions in the clean-correct cohort. In GEO it reduced injected-source selection by 28.75–47.06 points, but preserved only 67.97–82.35% of uninjected decisions. This substantial source-selection trade-off must accompany the defense result; reversing the correction worsened NQ2 accuracy.
Section 5.2, Controlling Susceptibility with a Low-Rank Update · Read the study
[c6] The authors explicitly restrict conclusions to controlled multiple-choice tasks and fixed-source GEO injection, leaving free-form generation and end-to-end retrieval/generation untested. They disclose generative-AI manuscript polishing and use of Meta-Llama-3-70B-Instruct to help construct anonymized controls and target-name-free appeals.
Section 6, Limitations; AI use statement · Read the study
[c7] In the selected decision heads on NQ2 and GEO, clean-to-persuaded decision-subspace patching recovered reference-answer probability comparably to or more than full-output patching; complementary-subspace effects were at most 2.2 percentage points in magnitude. This supports a causal role within this intervention system, not a claim about general answer accuracy. The reference is the correct option on NQ2 and the clean decision on GEO. Table 7 reports means with pointwise 95% percentile intervals from 2,000 paired question-cluster bootstrap resamples.
Appendix H, Subspace-restricted output patching; Table 7 · Read the study
[c8] Replacing each selected decision head's persuaded attention pattern with its clean pattern, without changing values, produced positive mean reference-probability recovery for every head on both datasets. NQ2 patches covered answer choices; GEO patches also covered documents. Table 8 uses 142–483 patched pairs per head on NQ2 and 231–528 on GEO, with 95% intervals from 2,000 paired question-cluster bootstrap resamples. These are full-vocabulary probability changes, not measured changes in generated-answer accuracy.
Appendix L, opening paragraph; Table 8 · Read the study
[c9] Feature-guided steering tested every A–D target on each held-out prompt, making 25% the exact target-selection baseline for target-independent outputs. Switch rates condition on the original decision being non-target; probability changes concern the target label token under the full-vocabulary distribution. This experiment must not be described as clean accuracy, persuasion success, or patching recovery. Evaluation used question-grouped rank-one test splits, clean and persuaded conditions, and a probability-margin filter of at least 0.10; Qwen's two heads were intervened on jointly.
Appendix P, Evaluation and Metrics; Table 11 · Read the study
[c10] Entity-anonymized controls were behaviorally selected separately for each model to preserve the clean-correct answer, rather than being an unselected anonymization sample. The pool contained 131 NQ2 questions correct under clean prompts for all four models, each with nine appeals. Only 3,548 of 4,716 potential model-specific records were accepted (75.2%); 892 lacked valid replacements and 276 failed the answer screen. Only 868 records were patching-eligible. The target name remained in the answer choices, and replacements could change prompt length.
Appendix Q, Behavioral selection; Control construction; Coverage and scope; Table 12 · Read the study
[c11] For selected upstream heads in the NQ2 persuasive-versus-anonymized comparison, target-answer-text queries sent 39.7% of full attention to the passage; entity pieces received 28.6% of that passage attention while occupying 5.5% of passage tokens on average. Reported mean casewise per-token enrichment was 6.1 times. These are descriptive attention destinations, not token-specific causal effects. Heads were selected using passage-attention and donor-directed patching-effect thresholds; aggregate statistics weight heads equally.
Appendix S, Causal identification and head selection; Attention destinations; Figure 27 caption · Read the study
[c12] Across five selected decision heads in four models on NQ2 and GEO, stronger upstream OV-map alignment generally occurred several layers before the decision heads, corresponding broadly with larger patching effects. The composition analysis characterizes geometric alignment; it does not independently establish causality. Its transport uses a corpus-averaged local Jacobian approximation estimated with frozen models from WikiText-103, using 7,813 sampled positions and 32 probes per position.
Appendix U, Results across models and datasets; Appendix T, Estimation procedure · Read the study
[c13] On held-out prompts, trained routing-key interventions increased NQ2 correct-answer accuracy under persuasion and reduced GEO injected-source selection across all four models. NQ2 accuracy rose from 53.70% to 75.13% for Llama, 67.61% to 79.91% for Gemma, 48.81% to 90.08% for OLMo, and 52.38% to 75.13% for Qwen. GEO injected-source selection fell from 83.66% to 36.60%, 51.63% to 22.88%, 47.06% to 1.96%, and 51.63% to 19.61%, respectively. These are intervention results, not evidence that GEO selection measures factual accuracy. Clean-decision agreement was 95.24–96.43% on NQ2 but only 67.97–82.35% on GEO. Reported uncertainty is question-level bootstrap 95% confidence intervals; NQ2 appeals remain grouped by question.
Appendix W, Test-set results, Table 17 and intervention note · Read the study
[c14] The intervention study selected model-specific NQ2 questions already answered correctly with a minimum 0.10 clean full-vocabulary probability margin, retaining all nine appeals without filtering for attack success. GEO used all 998 questions and the original model's clean choice as a pseudo-label, not ground truth. Splits were question-level. The rank-2 adapters were query-head-specific attention interventions, not globally modified checkpoints; test data were excluded from checkpoint selection. GEO explicitly optimized target suppression rather than correctness of an alternative, including when the target matched the clean reference.
Appendix W, Data and question-level splits; Intervention; GEO target-suppression objective; Checkpoint selection and controls · Read the study
[c15] NQ2 adapters trained and selected on two appeal families improved correct-answer accuracy on the excluded third family in all 12 model–family combinations: gains were 8.5–45.2 percentage points, with 89.3–97.9% clean-decision preservation. This transfer test held question splits fixed and excluded withheld-family validation and test results from selection. Estimates weighted questions equally and used 2,000 paired whole-question bootstrap resamples for 95% percentile intervals. Transfer was not uniformly equal to seen-family performance: Llama's emotion transfer gap was −10.7 percentage points, with interval [−20.6, −1.2].
Appendix W, Generalization to unseen appeal families, Tables 18–19 and concluding paragraph · Read the study
[c16] The persuasion-identified heads also contributed causally to evidence-supported answer changes in BBQ activation-patching tests. Llama and Gemma heads ranked first and restored concrete answers on 44.6% and 47.4% of 500 selected pairs per model; OLMo's identified head ranked second and restored 40.4%. Qwen heads ranked first and third but restored only 6.0% and 1.4%. These were model-specific cohorts selected for correct ambiguous and informative responses, with only 136 pairs shared across all four models. Probability gains were normalized over A/B/C, unlike the full-vocabulary persuasion measures; Table 21 uses 95% question-cluster bootstrap intervals.
Appendix Y, Model-specific cohorts; Intervention and recovery scores; Transfer of the identified heads; Table 21 · Read the study
[c17] Removing literal target names through local descriptive substitutions reduced, but did not eliminate, persuasion. Among question–appeal pairs answered correctly when clean, any-incorrect-answer rates fell from 46.65% to 33.59% for Llama (2,718 pairs), 34.02% to 21.47% for Gemma (2,925), 63.40% to 39.15% for OLMo (2,115), and 53.88% to 26.72% for Qwen (2,511). Table 22 uses 20,000 question-level bootstrap resamples for 95% intervals. Descriptions intentionally remained informative about the target. In separate confidence-filtered, successfully persuaded localization cohorts, the same identified heads led output and choice-list attention patching; those cohorts must not be confused with the behavioral denominators.
Appendix Z, Local edits and validation; Behavioral evaluation; Table 22; Localization cohort and interventions; Table 23 · Read the study
[c18] The authors explicitly limit deployment implications: controlled multiple-choice findings do not establish effectiveness in open-ended generation or real-world interactive systems. They also disclose dual-use risk, including assistance to persuasive prompting, retrieval poisoning, and GEO-style manipulation, and frame the work as defensive rather than an attack recipe.
Appendix AB, Broader Impacts · Read the study