The practical brief
The finding
The diagnostic agent recovered a deliberately altered catalog signal in 39 of 100 synthetic trials; a small Gemini pilot also demonstrated live visibility measurement. [c1] [c2] [c3] [c4] [c5] [c6] [c7]
Why it matters
Retailers could use this approach to investigate recommendation gaps, but neither its visibility shares nor its diagnostic rankings establish which catalog changes will improve recommendations or sales.
| Evaluation | Reported result and scope |
|---|---|
| Overall synthetic signal recovery | 39/100 trials (39.0%); Wilson 95% CI: 30.0%–48.8% [c2] |
| Top-three-correlate subset | 39/61 qualifying synthetic trials (63.9%); the same 39 successes [c2] |
| Live extraction coverage | 21/30 Gemini responses (70.0%); 10 queries repeated three times [c3] |
| Catalog intervention effectiveness | Qualitative interpretation: not established; controlled intervention testing remains proposed [c4] [c7] |
What to try
BLURSOR’s practical interpretation
Our interpretation: use the output as a human-reviewed investigation backlog, report extraction coverage alongside visibility share, and test proposed changes before scaling them.
Study boundary
Most evidence is synthetic. Live testing covered one platform and 30 calls, extraction lacks human validation, and no completed catalog intervention established uplift.
Do not turn a visibility gap into an automatic catalog fix
Before spending on product-page changes to improve AI discoverability, separate two questions: does an assistant recommend your retailer, and would changing your catalog make it recommend you more? This study offers a prototype for the first question and association-based suggestions for the second—not proof that those suggestions work.
The running-shoes evaluation used 23 SKUs—individual catalog products—across seven retailers and 34 discovery queries. Most testing used simulated platform responses and synthetic catalog signals. The system extracts recommendations, measures retailer visibility, and asks a diagnostic agent to rank attributes associated with recommendation gaps, such as reviews, product information, or pricing.
Its live check was much smaller: 10 generic queries sent to Gemini gemini-2.5-flash three times each, producing 30 responses. This is feasibility evidence for a measurement workflow, not a validated merchandising service.
The score measures retailer attribution, not demand
Agentic Share-of-Search is a rank-weighted share of extracted recommendations attributed to tracked retailers. Earlier mentions receive more weight. It measures where the assistant says someone can buy, rather than catalog presence, customer preference, or sales.
Its denominator matters: only extracted mentions that resolve to retailers in the competitive set count. Responses without extractable mentions contribute nothing. Duplicate product mentions count once at their earliest rank; mentions attributed to multiple retailers split their weight equally.
In the live pilot, extraction succeeded for 21 of 30 responses, or 70.0%. Zappos received 39.2% of the measured retailer share—not 39.2% of shoppers or all answer content. The paper attributes the leaderboard shift from simulated data to Gemini naming marketplaces even for products from brands that also sell directly.
- Read share together with extraction coverage: omitted answers can change what the leaderboard represents.
- Distinguish recognition of a product brand from attribution to the retailer selling it.
The diagnostic test recovered signals, not profitable actions
The main test deliberately degraded one catalog signal in each of 100 synthetic trials while leaving the original visibility outcomes unchanged. Success meant the agent’s first recommendation named that altered signal. It succeeded in 39 trials: 39.0%, with a 95% confidence interval of 30.0%–48.8%. Uniform random selection among 14 signals had an expected success rate of 7.1%.
The higher 63.9% figure applies only to 61 trials where the altered signal ranked among the three strongest correlations. It reuses the same 39 successes; it is not overall accuracy or evidence that recommendations improve visibility.
A correlation is an association, not evidence that changing one attribute causes an outcome. Retailer identity, brand, price tier, and product quality could explain the patterns. The agent’s priority score therefore orders a possible work backlog; it does not estimate expected uplift.
Use the prototype to frame tests, not promise returns
Our practical interpretation is to treat these diagnostics as hypotheses for human review. Inspect underlying answers and retailer assignments before committing resources, particularly because extraction and attribution have not been checked against human coding.
The author proposes—but has not completed—a controlled study: apply selected changes to a treatment group, retain a matched within-retailer control, and measure visibility weekly for four weeks. That distinction is essential: even replacing synthetic signals with real catalog data would not, by itself, establish causality.
The study cannot tell an owner which change will increase sales, whether the findings transfer to another category, or whether results persist across platforms and time. The author, Spandan Ghose Chowdhury, lists Georgia Institute of Technology’s College of Computing as his affiliation and discloses Claude assistance with code scaffolding, drafting, and editorial review. The supplied paper contains no funding or commercial-interest declaration.
The source and evidence
Study limitations and disclosures
- Most testing used simulated responses and synthetic catalog signals. Live evidence covers only Gemini, 10 queries, and 30 calls; the study is limited to one category and one time period.
- Synthetic signal recovery retained the original visibility outcomes. It does not establish the best managerial intervention, causal visibility uplift, or sales effects; correlations may reflect retailer, brand, price-tier, or product-quality differences.
- Extraction and retailer attribution have not been validated against human coding. Visibility shares exclude responses without extractable mentions and must be read alongside extraction coverage.
- The author lists Georgia Institute of Technology’s College of Computing as his affiliation and discloses Anthropic Claude assistance for code scaffolding, manuscript drafting, and editorial review, with author review and responsibility. No funding or commercial-interest declaration appears in the supplied paper.
Evidence behind this briefing
[c1] ASoS measures rank-weighted retailer attribution in AI recommendations, not catalog presence or sales. Its denominator includes only extracted mentions resolving to tracked retailers; extraction failures are excluded, duplicate SKU mentions use their earliest rank, and multi-retailer mentions split their weight equally.
Page 4, Methodology — Stage 2, Metric specification · Read the study
[c2] In 100 synthetic signal-recovery trials, the diagnostic agent recovered the ablated signal in 39/100 trials: 39.0%, Wilson 95% CI 30.0%–48.8%. The reported 5.5× baseline comparison uses expected approximately 7% success. The 63.9% result is conditional on 61 top-three-correlate trials and reuses the same 39 successes, rather than showing overall accuracy or intervention effectiveness.
Page 7, Empirical Results — Diagnostic-agent reliability via signal ablation · Read the study
[c3] The live test used Gemini gemini-2.5-flash only: 10 generic queries repeated three times, totaling 30 calls. Extraction succeeded in 21/30 responses (70.0%); Zappos received 39.2% of the live retailer share. The paper attributes the difference from mock data to marketplace attribution, not demonstrated consumer preference or purchasing.
Page 9, Empirical Results — Pipeline robustness and live-platform validation · Read the study
[c4] The diagnostic rankings are cross-sectional associations, potentially confounded by retailer, brand, price tier, or product quality. Priority scores order a merchandising backlog; they do not estimate visibility uplift, establish predictive performance, or demonstrate causal effects from changing catalog signals.
Page 6, Methodology — Estimand, unit of analysis, and the association–prediction–intervention distinction · Read the study
[c5] Most analysis uses synthetic mock responses and a synthetic signal database. Generalization to real retail catalogs or multiple AI platforms remains untested; the live evidence is limited to Gemini, 10 queries, and 30 calls.
Page 9, Discussion — Boundary conditions and limitations · Read the study
[c6] Extraction and retailer attribution have not been validated against human coding. The main study covers only one category, one live platform, and one time period, so its machine-coded visibility figures and diagnostic outputs should be treated as prototype evidence.
Page 10, Discussion — Boundary conditions and limitations · Read the study
[c7] A proposed, not completed, intervention study would apply selected catalog changes to a treatment group, retain a matched within-retailer control, and measure weekly ASoS for four weeks with effect sizes and confidence intervals. This future design is intended to supply causal evidence missing from the ablation evaluation.
Page 10, Discussion — Future work: proposed experimental design for intervention closure · Read the study
[c8] The author discloses using Anthropic Claude for code scaffolding, manuscript drafting, and editorial review, with author review and responsibility for the publication.
Page 11, Declaration on GenAI Usage · Read the study