The practical brief
The finding
In controlled e-commerce tests, one AI ranker promoted unsupported-rich product descriptions over clean descriptions; susceptibility was weaker or absent in two other models. [c2] [c4] [c7] [c12]
Why it matters
A prominent product recommendation is not necessarily a signal that its claims have been verified.
| Target variant | No evidence penalty | Construction-label evidence penalty |
|---|---|---|
| Unsupported-rich | 65% of unsupported-rich ranking cases placed in the top three | 43% of unsupported-rich ranking cases placed in the top three [c3] |
| Laundering | 61% of laundering ranking cases placed in the top three | 39% of laundering ranking cases placed in the top three [c3] |
What to try
BLURSOR’s practical interpretation
Our interpretation: assess AI ranking performance and claim substantiation separately, and ask evidence-checking vendors how they handle missing support and disputed labels.
Study boundary
These synthetic ranking experiments do not establish effects on deployed AI visibility, citations, customer trust or sales, and the automatic evidence judges failed the study’s reliability gate.
Do not treat ranking position as verification
If you are deciding whether to invest in richer product descriptions or an AI evidence-checking service, separate two questions: does the text attract a ranker, and can its claims be substantiated? This study shows those outcomes can diverge. It does not show that adding unsupported claims improves visibility in live AI search.
Researchers tested 50 English shopping queries derived from the ESCI dataset, producing 1,950 controlled ranking cases. They compared clean descriptions with richer supported, unsupported and neutral variants, plus an attack variant called laundering. Rich variants were matched for format and volume; clean descriptions were not length-matched.
The study also changed evidence packets while keeping descriptions fixed. These packets supplied material against which claims could be checked. Here, “unsupported” means unsupported by the supplied packet—not necessarily false in the wider world. No real listings were changed or live rankings targeted.
Fabricated richness gained ground on one ranker
On Qwen2.5-7B, unsupported-rich descriptions gained +0.065 to +0.092 in normalized rank relative to clean descriptions across three claim profiles. Normalized rank is a scaled position measure, not accuracy or conversion. These gains survived the study’s correction for multiple statistical tests; supported-rich and neutral controls did not significantly separate from clean.
The effect was model-dependent: evidence was weaker on MiMo-v2.5 and absent on GLM-5.3-Flash. These were within-model tests, not proof of formal differences between models.
A separate defense experiment used a frozen DeepSeek-Flash scorer and penalized query-relevant claims lacking packet support. With construction-time labels—labels known from how the benchmark was built—the unsupported-rich top-three rate fell from 0.65 to 0.43 at penalty weight 40. Synthetic attestations restored the original rate without changing the description. That isolates evidence availability; it does not validate authentication or an automatic defense.
Evidence checks need coverage as well as accuracy
The realistic business interpretation is not “write less detail.” It is that attractive detail and verified detail are different properties. An evidence-based system can also penalize supported descriptions when its evidence is missing: withdrawing the tested attestations increased supported-rich false suppression by 0.307 with text unchanged.
These are practical evaluation questions, not tactics proven to improve discoverability:
- For important product claims, document the supporting material and its relationship to the specific product; do not assume richer wording demonstrates verification.
- Ask an evidence-checking vendor to distinguish missing evidence from contradictory evidence, and explain how disputed assessments can be reviewed.
- Evaluate ranking gains separately from evidence accuracy and inappropriate demotions. One favorable metric cannot establish all three.
The automatic defense is not ready-certified
All three automatic judges failed the preregistered reliability gate against 370 human-labeled claims from just 10 queries. The best local judge scored macro-F1 0.374 and kappa 0.442—classification performance and agreement measures—below required thresholds of 0.75 and 0.60.
The human reference itself has unresolved disagreements: five clean identity claims account for measured collateral, while roughly 10% of attack claims were labeled attested. Construction labels are not human gold, and zero suppression of protected variants in some experiments does not settle these disputes.
Other limits matter: only the target had audited claims, so competitors could not be penalized; evidence coverage was uneven and its reported split remains unreconciled; penalty weight 40 was not a validated deployment optimum. Displayed top-three rates do not establish the separately defined promotion-reduction, utility or grounding outcomes. Transfer to local services, B2B and web passages remains untested, as do real-world evidence omission rates.
The source and evidence
GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings
Study limitations and disclosures
- Closed-world labels measure support in supplied packets, not truth: world-true claims without packet evidence are unsupported, and unverifiable claims remain abstentions (Sections 3.2–3.4). Open-Web discovery, entity resolution and universal fact checking are outside scope (Section 4.7).
- The independent statistical unit is the query (n=50), with three order seeds averaged within query. Rich arms are length-matched, but clean is not. MiMo has only one significant corrected contrast against neutral-matched; GLM has none. These are within-model tests, not formal between-model differences or equivalence tests (Sections 3.4–3.5 and 6.1).
- Defense experiments use a separate pointwise DeepSeek-Flash scorer, not the phenomenon listwise ranking orders. Reported top-3 success-rate drops do not by themselves establish the frozen clean-referenced Promotion reduction endpoint (Sections 3.6, 5.3 and 5.5).
- Construction labels are oracle by provenance, not human gold. Against human gold they show 9.3% claim-level false penalties. The human-gold ranking arm has maximum protected-arm collateral of 0.300 on 390 holdout cases (Sections 6.5–6.6).
- Automatic-label results are diagnostic because the gate failed. Agreement metrics are point estimates from one 10-query holdout without interval estimates. The API judge ran only on the holdout; the two local judges share a model family and are not independent evidence (Sections 6.5–6.6).
- Packets include listing assertions, buyer-review atoms covering 42/50 ASINs, and 13 registry atoms across 18 credential queries; one registry case is contested. Twin attestations are synthetic construction-tier additions, not authoritative independent evidence (Sections 3.3 and 5.2–5.4).
- The tuning rule was preregistered for this revision, not earlier drafts. Oracle and L2 selections reach the tested grid edge at lambda 160; the main lambda 40 comparisons are not development-selected optima (Sections 5.6 and 6.9).
- Label-noise ranges are deterministic simulation ranges, not confidence intervals or repeated-judge variability. Figures depicting products and packet operations include schematic illustrations, not measured product examples (Figures 1–2 and 6).
- The two local judges belong to one model family; the API judge runs only on the holdout. Gate failures do not establish failure across judge families.
- All defense endpoints use the same pointwise base scorer, DeepSeek-Flash; rates and relative gains are not established for other scorers.
- Useful automatic-label ranking effects do not validate an audit ledger; no alternative endpoint-based admission threshold was validated.
- Aggregate API-judge errors do not establish that plausibility or apparent attestation caused the errors.
- Construction labels assign no independent_source_supported claims, whereas humans assign 13: three unsupported-rich, five neutral, and five laundering claims.
- Evidence thinning suggests a mechanism of harm to publishers with thin coverage, not its real-world distribution. Retrieval repair, packet-omission appeals, and proposed governance measures remain unvalidated.
- Two annotators independently labeled each of 370 claims; automated suggestions are disclosed separately, identities remain private, and annotation contained no personal information. ESCI and Amazon Reviews 2023 are used under research terms.
Evidence behind this briefing
[c1] The benchmark uses 50 English ESCI-derived e-commerce queries, 13 target variants per query and 1,950 ranking cases. Only targets carry audited claims, so the defense can lower targets but cannot reorder distractors.
Section 5.2, Benchmark and Candidate Arms · Read the study
[c2] Qwen2.5-7B rewarded unsupported-rich descriptions over clean descriptions across three claim profiles: paired normalized rank gains were +0.065 to +0.092, with Holm-adjusted p values from 0.035 to 0.0011. This is rank promotion, not factual accuracy; susceptibility was not universal across tested models.
Section 6.1, Unsupported-Specific Rank Promotion; Table 3 · Read the study
[c3] On the frozen pointwise scorer, construction-label reranking at lambda 40 reduced unsupported-rich top-3 rates from 0.65 to 0.43 and laundering from 0.61 to 0.39, with zero protected-arm suppression by construction. Fixed-text synthetic packet twins restored baseline rates; this probes attestation availability, not real-world authentication or validated automatic defense.
Sections 6.3–6.4, Table 5 and Packet Twins · Read the study
[c4] All three automatic evidence judges failed the preregistered human-gold reliability gate. The best local judge had macro-F1 0.374 and kappa 0.442 versus thresholds 0.75 and 0.60, on 370 claims from a 10-query holdout; successful parsing is not evidence-label reliability.
Section 6.5, The Automation Gap; Tables 6–7 · Read the study
[c5] In the controlled packet-thinning intervention at lambda 40 over 1,950 cases, fully removing the tested attestations increased supported-rich false suppression by 0.307 despite unchanged text. This measures engineered evidence loss, not its prevalence in deployed search.
Section 6.7, Coverage Dose–Response; Table 10 · Read the study
[c6] The adverse human-gold reranking result remains unresolved: five clean identity claims labeled unsupported account for measured collateral, while roughly 10% of attack claims labeled attested contribute to weaker suppression. The authors do not resolve whether this reflects gold noise or a definitional disagreement.
Section 6.6, Pending human-gold interpretation; Table 9 · Read the study
[c7] Unsupported-specific promotion was significant on Qwen2.5-7B, but weaker or absent on the other two tested models; transfer beyond ESCI e-commerce remains untested.
Section 8, Ranker and domain scope · Read the study
[c8] Support is measured against supplied evidence rather than open-world truth. Coverage is uneven and its split unresolved; all 50 queries remain in headline estimates, while coverage conclusions rely on within-query thinning rather than comparisons with uncovered queries.
Section 8, Supplied evidence and incomplete coverage · Read the study
[c9] Human-gold estimates use 370 claims from only 10 queries. Clean-arm collateral and weaker attack suppression remain unresolved results, not evidence that human judgment itself causes harm.
Section 8, Human-gold size and unresolved judgments · Read the study
[c10] The development sweep does not identify a deployable optimum; main comparisons retain lambda=40 and are not a tuned-and-frozen holdout evaluation.
Discussion immediately following Table 12 · Read the study
[c11] Only the target is penalized because distractors have no claims. Displayed raw top-3 rates and Success reduction do not establish the distinct primary Promotion reduction endpoint or the defined utility and grounding measures.
Section 8, Target-only intervention and endpoint scope · Read the study
[c12] The attacks used synthetic descriptions in a controlled benchmark without altering real content or targeting live rankings; the authors disclose dual-use risk and release evaluation apparatus.
Section 9, Ethical Considerations · Read the study