How AI decides what to cite, rank, and surface

Catch up with the papers shaping AI visibility. Each distill explains the findings, supporting evidence, and practical implications in plain English.

How a distill is made

  1. We find papers that address AI visibility, ranking, retrieval, and citation behavior.

  2. We examine the methods, results, and limitations, then isolate findings with practical implications.

  3. We publish the distill in plain English, with the supporting evidence and a direct link to the paper.

Recent Digests 42 articles
01
Sep 1, 2026 9 min read

Query language decides which market a generative search tool can recommend

Your prompt’s language can pick the market—IP only swaps the brands.

arXiv:2608.30052
02
Sep 1, 2026 9 min read

Rank-only incentives can make content look better to a ranker while drifting away from human quality

When “rank-only” incentives silently degrade what humans call good

arXiv:2608.30466
03
Aug 25, 2026 9 min read

RAG systems can collapse when they start citing their own output

Self-authored retrieval can “lock” RAG into collapsed answers—at scale.

arXiv:2608.22118
04
Aug 20, 2026 8 min read

AI search mode cuts publisher referrals without making search feel better

AI summaries shift clicks away from publishers and toward competitor engines—at the cost of user trust

arXiv:2608.18352
05
Aug 18, 2026 9 min read

Generative search can erode the web it depends on

How extraction siphons the web until it can’t renew itself

arXiv:2608.15896
06
Aug 10, 2026 9 min read

AI assistants miss 85.6% of venues in a complete Bali market census

Most venues are missing from AI answers—what a true census audit reveals

arXiv:2608.07069
07
Jul 28, 2026 9 min read

Citation type predicts who grounded LLMs name — and roster-based visibility misses almost all of it

Where grounded LLMs actually start naming people—and why most “visibility” measurements miss

arXiv:2607.23893
08
Jul 21, 2026 10 min read

DRNoise shows how one plausible false document can knock deep research agents off course

One misleading “direct claim” can derail an agent that’s otherwise correct

arXiv:2607.17291
09
Jul 20, 2026 9 min read

Chinese generative search cites only a small slice of available brand sources, and external quality scores do not predict what surfaces

Why citation coverage is sparse—and what actually predicts which sources get surfaced

arXiv:2607.15771
10
Jul 16, 2026 8 min read

GEO can change citations inside a fixed context, but it doesn’t show durable organic visibility

Why GEO gains don’t translate into durable visibility (and which levers actually hold).

arXiv:2607.14035
11
Jul 16, 2026 9 min read

LLM brand answers are mostly unstable because language changes the signal

Variance components explain why brand answers won’t stabilize—and what to sample instead

arXiv:2607.13304
12
Jul 7, 2026 9 min read

Open-web search answers more questions, but it makes source trust much harder to control

When open-web search boosts coverage, it quietly degrades source trust

arXiv:2607.05217
13
Jun 25, 2026 9 min read

LLMs mostly source “brand reputation” from other people’s pages

What LLMs treat as “brand reputation” is mostly other people’s web pages, not the brands themselves

arXiv:2606.25787
14
Jun 23, 2026 10 min read

AI brand “ownership” is moderately concentrated, but the winner changes by model

Who gets the “top pick” in AI recommendations—and how consistent is it across models?

arXiv:2606.23057
15
Jun 23, 2026 10 min read

English-only AI reputation monitoring misses local champions in multilingual markets

English-language prompts create a measurable “local-visibility” blind spot across languages

arXiv:2606.23165
16
Jun 23, 2026 9 min read

AI visibility breaks down by entity, not just by mention count

Why “mention” counts fail: fabricated citations scale differently by entity and query context

arXiv:2606.21595
17
Jun 19, 2026 9 min read

AI search visibility starts with brand stature, not prompt tweaks

Why the first-run visibility gap between big brands and everyone else is so persistent

arXiv:2606.20065
18
Jun 17, 2026 9 min read

Incumbent brands get a built-in advantage in LLM recommendations — but only until a competitor clears a narrow threshold

How LLM recommenders lock in incumbent brands—and the small tweaks that break it

arXiv:2606.17443
19
Jun 16, 2026 10 min read

LLM search agents can be pushed to endorse manipulated web content

When “search-and-answer” becomes endorsement for sale

arXiv:2606.16821
20
Jun 12, 2026 9 min read

One Polluted Page Is Enough to Hijack LLM Recommendations

Why one poisoned search result is enough to hijack LLM recommendations

arXiv:2606.13610
21
Jun 10, 2026 10 min read

AI brand recommendations move people onto the open web through search, not just mention counts

When an assistant says the brand name: what actually moves browsing

arXiv:2606.10907
22
Jun 9, 2026 9 min read

Safety-trained RAG models can turn a prompt injection into brand suppression

Why safety alignment can turn retrieval-time injections into brand-level anti-promoters

arXiv:2606.09204
23
Jun 8, 2026 8 min read

FullCite gets better quote grounding by separating the document from the evidence span

FullCite turns inline citation into a document-plus-span problem, sharply improving quote-level grounding on ASQA.

arXiv:2606.07130
24
Jun 4, 2026 10 min read

ChatGPT referral spikes can overstate AEO unless you control for platform growth

Why “2x on ChatGPT” stories can be misleading without a tailwind control

arXiv:2606.04362
25
Jun 2, 2026 7 min read

LLMs Score 94% on Cultural Knowledge Tests and 40% When the Answer Choices Are Removed

When Cultural Knowledge Doesn't Transfer to Cultural Reasoning

arXiv:2606.01879
26
Jun 2, 2026 7 min read

English Prompts Suppress Bengali Cultural Knowledge Even When Local Evidence Is Provided

How Prompt Language Rewrites Cultural Knowledge Before the Model Even Answers

arXiv:2605.30481
27
Jun 2, 2026 7 min read

LLM Fact-Checkers Score Well But Retrieve the Wrong Sources

Where LLM Fact-Checkers Go Wrong on Sources

arXiv:2605.30241
28
Jun 2, 2026 7 min read

An LLM Agent That Learns From Its Own Mistakes Beats Human Fact-Checkers on Health Misinformation 89% of the Time

When the Agent Learns From Its Own Corrections

arXiv:2606.02215
29
Jun 1, 2026 7 min read

AI Overviews Sent Users to Reddit. AI Mode Is Taking Them Back.

Google AI Overviews drove a 12% rise in Reddit engagement, but AI Mode reversed those gains for experiential communities by substituting conversation for human discussion.

arXiv:2605.16428
30
Jun 1, 2026 7 min read

Ecosystem GEO Beats Page-Level Optimization by Up to 31 Points for Agent Search

Coordinating a multi-page evidence ecosystem raises LLM search agent recommendation rates by up to 31 percentage points over the best single-page GEO baseline.

arXiv:2605.12887
31
Jun 1, 2026 7 min read

When AI Cites AI: The Synthetic Source Problem in Generative Search

An audit of ChatGPT, Copilot, Gemini, and Perplexity finds ~16% of cited sources are AI-generated — with Copilot citing synthetic content in nearly 3 of every 10 citations.

arXiv:2605.23684
32
Jun 1, 2026 6 min read

Frontier LLMs Hallucinate Up to 38% of Scientific Citations

Six frontier LLMs hallucinate 12–38% of scientific citations; a new agentic retrieval system hits zero hallucination at 30% better F1 and $0.05 per query.

arXiv:2605.14306
33
Jun 1, 2026 6 min read

Rewording a Buying Question Changes the Brands AI Recommends More Than Switching Models Does

Cosmetic prompt rewording drops AI brand-recommendation overlap by 21–32 percentage points — more divergence than switching providers entirely, across 12,000 runs.

arXiv:2605.27440
34
Jun 1, 2026 6 min read

Query-Specific Expiry: Why 'Recent' Isn't the Same as 'Fresh'

Baidu's Aurora-Expiry uses RAG-augmented LLMs to infer query-specific expiration thresholds, cutting median document age 12.81% for time-sensitive queries in a 14-day live A/B test.

arXiv:2605.13052
35
Jun 1, 2026 8 min read

RAG Doesn't Flatten the Brand Hierarchy — It Just Moves Where You Lose

A 37,000-run audit of 533 brands finds RAG preserves the brand hierarchy: L4–L5 specialists face 48–52% invisibility while L1 leaders surface universally but convert at only 25–41%.

arXiv:2605.27439
36
May 31, 2026 7 min read

Semantic Metadata Makes Agents More Reliable, Not Smarter

Schema.org markup gives retrieval agents 65.7% higher FAIR-compliant precision — but cuts query coverage by 29% where publishers haven't adopted it.

arXiv:2605.28787
37
May 26, 2026 6 min read

One Bad Search Result Breaks Frontier AI Agents — Completely

A Microsoft study shows a single false top search result drops GPT-5 accuracy from 65% to 18% — while humans solve the same queries at 93% — exposing a critical gap in agentic RAG deployments.

arXiv:2603.00801
38
May 22, 2026 8 min read

AI Research Agents Fail in Ways Their Own Tests Can't See

A new SoK survey of 118 works shows that agentic RAG's iterative retrieval and memory systems introduce failure modes that static metrics and current benchmarks cannot detect.

arXiv:2603.07379
39
May 22, 2026 7 min read

Most Pages Get Zero AI Citations. Editing 5% Won 40% More.

A new paper from Virginia Tech maps four failure modes that prevent pages from being cited in AI-generated responses. 43% of relevant pages receive zero citations under baseline conditions.

arXiv:2603.09296
40
May 22, 2026 7 min read

AI Cites Sources It Never Checked — and Half Don't Hold Up

No LLM verifies even half its citations under any tested condition — and adding temporal cutoffs or other deployment constraints collapses verifiability to near zero.

arXiv:2603.07287
41
May 22, 2026 8 min read

Generative Search Citation Share Is a Noisy Estimator, Not a Score

A new statistical framework shows that single-run citation share metrics from Perplexity, SearchGPT, and Gemini carry confidence intervals wide enough to make most apparent SEO gains statistically indistinguishable from noise.

arXiv:2603.08924
42
May 22, 2026 7 min read

Schema Markup Barely Helps AI Find You. Restructuring Does.

Rewriting pages as entity documents lifted AI answer accuracy ~30%; adding JSON-LD did almost nothing — it gets cut before indexing.

arXiv:2603.10700