The practical brief
The finding
The paper argues that GEO should be evaluated across a conversation trajectory, not as a one-answer lift, because an early citation can increase later citation odds through what the authors call conversational capture. [c1] [c2] [c5] [c6] [c7] [c8]
Why it matters
If you choose between GEO tactics from single-turn results alone, you may mis-rank them under the paper’s model. In its worked example at 10 turns, the modeled feedback omitted by single-turn testing was larger than the direct effect, and single-turn versus trajectory rankings agreed only weakly.
| Measure | Single-turn view | Trajectory / multi-turn view |
|---|---|---|
| Feedback term | 0 by construction | 1.48 at T=10 [c5] [c6] |
| Direct term | 1.20 at T=10 by naive T·L1 extrapolation | 1.20 at T=10 [c6] |
| Total modeled gain | 1.20 at T=10 if you extrapolate from one turn | 2.68 at T=10 [c6] |
| Method ranking agreement | Ranks methods by L1 only | Kendall's tau = 0.4 between single-turn and trajectory rankings [c7] |
What to try
BLURSOR’s practical interpretation
Our interpretation: keep single-prompt citation tests for screening, but make final evaluation multi-turn when customer journeys usually involve follow-up questions. Track cumulative conversational visibility across turns, measure machine-side persistence on identical query sequences first, and label any human-side estimate as provisional unless you can support the required counterfactual with observed logs or a validated simulator.
Study boundary
This is a theory-and-modeling paper from HAI ’26, not a deployed-engine field study. Its numbers are model-derived, the sign of feedback is an empirical question rather than an assumed benefit, M2 is a nested residual rather than a pure standalone human effect, and the authors also note imperfect LLM-as-judge scoring, no real trust or perception measurement, open-source RAG versus closed-engine gaps, and a simplified two-source competition model.
The decision risk is using too small a test
If you are deciding whether to fund a GEO rewrite, vendor, or internal experiment, this paper says the usual test unit may be wrong. A one-prompt benchmark treats each query as independent, but in a real AI conversation the answer can shape the next question and the next retrieval step.
The paper’s central idea is conversational capture: once a source is cited early, its chance of being cited again can rise later, even if the user’s need has narrowed or drifted. For a business owner, that means the business value of AI visibility may depend less on one answer and more on whether your source keeps resurfacing across a session.
- Single-turn GEO testing assumes queries are independent of the system.
- Multi-turn dialogue breaks that assumption because answers influence later turns.
- Early citation can create persistence, not just a one-time mention.
What the framework measures instead
This paper does not propose a new GEO tactic. It proposes a new evaluation frame: measure visibility over the whole conversation path. Its main trajectory metric is cumulative conversational visibility, meaning the total visibility a source accrues across turns, with optional weights if earlier or later turns matter more.
The authors then split trajectory gain into two parts. The direct term is the summed GEO lift with the conversation’s mediator path held to its no-GEO realization. The feedback term is the extra visibility created only because GEO changed later history or later questions.
The framework further nests that feedback split into machine-side and human-side channels. Machine-side capture means the engine is more likely to resurface something it already cited because it conditions on dialogue history. Human-side capture means earlier answers nudge follow-up questions toward that source. Under the main nested attribution, M1 is directly identifiable on identical queries, while M2 is the residual after M1, so it also contains interaction between machine and human effects rather than a perfectly isolated human-only quantity.
- Trajectory metric: cumulative conversational visibility across turns.
- Direct term: lift summed while the mediator path stays at its no-GEO realization.
- Feedback term: extra gain created by the closed loop.
- M1 can be measured from system comparisons; M2 needs a counterfactual branch.
Why one-turn rankings can fail in the model
The paper first makes a structural point: in single-turn evaluation, the feedback term is zero by construction. So a one-turn test cannot observe the compounding effect the framework is designed to capture.
In the worked example at T=10, modeled trajectory gain was 2.68, split into a direct term of 1.20 and a feedback term of 1.48. Within that illustration, the omitted feedback was larger than the direct effect. The figure in the paper further splits that modeled feedback into 0.85 machine-side and 0.63 human-side under an illustrative parameter split.
That changed method choice inside the model. Across five stylized methods, single-turn and trajectory rankings had Kendall’s tau of 0.4, and the single-turn winner ranked only third by trajectory gain. The paper also says disagreement appears often across many random method draws in the same model, especially when immediate salience and capture strength move in opposite directions. Still, the authors explicitly limit this to the model: whether real GEO methods are separated enough to create the same misranking is an empirical question.
- Single-turn evaluation forces feedback to zero.
- Worked example at T=10: feedback 1.48 exceeded direct gain 1.20.
- Misranking is shown within the model, not demonstrated on commercial engines.
What this paper cannot tell you yet
The most important limit is external validity. This HAI ’26 proceedings paper with DOI 10.1145/3841580.3841626 is theory and modeling, not a live study of ChatGPT, Gemini, AI Overviews, or Perplexity. The authors collected no human annotations, and the numeric examples demonstrate what the model permits rather than estimating real-world effect sizes.
The paper also does not assume feedback is always positive. Its main examples use a confirmatory reinforcement regime, where earlier citation helps later citation. But the authors explicitly note a corrective regime too: follow-up questions could steer users toward competitors, making the human-side channel negative and causing single-turn evaluation to overestimate GEO instead of underestimate it.
There are also measurement and scope limits relevant to marketers. The authors say LLM-as-judge visibility scoring is imperfect; their proposed empirical path uses open-source RAG rather than closed commercial engines; they measure no real user trust or perception; and their main urn model reduces many competing sources to one aggregate competitor, leaving multi-optimized-source competition outside scope. The paper discloses one author affiliation with UniConvo Inc. and JSPS KAKENHI Grant Number 25K21201 funding.
- No deployed-engine measurement or real-user trust study.
- Bias direction is empirical: single-turn can understate or overstate GEO.
- M2 is not a clean standalone human estimate under the main nested decomposition.
- Open-source RAG validation would still need replication across models and engines.
Our practical reading for marketers
Our interpretation, not an experimentally proven tactic: change the evaluation stack before increasing GEO spend. A one-prompt citation test is still useful for triage, but it is too narrow as the final decision tool when buyers usually ask follow-up questions.
A realistic next step is to build short multi-turn scripts around actual research or purchase journeys. Track cumulative visibility across turns, note whether first-turn citation persists, and measure machine-side persistence first by comparing history-conditioned versus stateless or reset conditions on the same query sequence. Treat any estimate of human-side drift more cautiously, because the paper says that counterfactual requires a simulated or otherwise reconstructed branch.
If you report results internally, separate what is observed from what is inferred. Observed: per-turn visibility and persistence across scripted follow-ups. Inferred: how much of that came from user steering rather than engine memory. That distinction is the most operationally useful part of this paper.
- Use single-turn tests for screening, not sole budget decisions.
- Add multi-turn scripts based on real follow-up behavior.
- Prioritize measuring M1 first because it is the identifiable channel.
- Label M2 and total behavioral interpretation as more provisional.
The source and its limits
Findings, methods, and limitations from the full paper relevant to marketers deciding how GEO should be evaluated in multi-turn AI conversations.
- The paper is theoretical and model-based rather than an empirical study with deployed systems or real users: “This paper contributes theory and modeling rather than a new human study. We collect no human annotations.” (§1, Contributions)
- The computational illustration is not an estimate of real-world effect sizes: “It demonstrates the constructs rather than providing evidence about deployed engines.” (§7, first paragraph)
- Human-side estimates would require counterfactual user modeling in future empirical work: “Only this counterfactual requires a simulated branch.” (§10, Dependence on user simulation)
- The authors note measurement limitations in future validation, including imperfect visibility scoring and lack of real trust/perception measures: “(i) LLM-as-judge visibility scoring is imperfect” and “(iii) We measure no real human trust or perception; that is future work.” (§10, Other limitations)
Evidence behind this briefing
[c1] The paper’s main thesis is that single-turn GEO metrics misestimate impact because they ignore multi-turn feedback.
§1, Thesis · Read the study
[c2] The paper defines conversational capture as persistence: early citation raises later citation likelihood, potentially apart from current relevance.
§2.3, “Conversational capture” · Read the study
[c3] The core trajectory-level visibility metric is cumulative conversational visibility across turns, optionally weighted by attention or salience.
§5.1, Eq. (2) · Read the study
[c4] The framework decomposes trajectory gain into direct per-turn lift and additional feedback caused by GEO-shifted queries and history.
§5.2, Eq. (5) · Read the study
[c5] Single-turn evaluation omits the feedback term by construction; it appears only when history or a human closes the loop.
§5.2, Key property · Read the study
[c6] In the worked example at T=10, modeled feedback exceeds direct gain, so single-turn extrapolation understates modeled conversational payoff.
§7.2, paragraph before Table 1 · Read the study
[c7] In the five-method illustration at T=10, single-turn and trajectory rankings disagree, so optimizing the single-turn metric can pick the wrong method within the model.
§7.3, paragraph after Table 2 · Read the study
[c8] The quantitative illustration is model-derived, not evidence from deployed engines or real users.
§1, Contributions, C5 · Read the study