Catch up with the papers shaping AI visibility. Each distill explains the findings, supporting evidence, and practical implications in plain English.
We find papers that address AI visibility, ranking, retrieval, and citation behavior.
We examine the methods, results, and limitations, then isolate findings with practical implications.
We publish the distill in plain English, with the supporting evidence and a direct link to the paper.
Your prompt’s language can pick the market—IP only swaps the brands.
arXiv:2608.30052When “rank-only” incentives silently degrade what humans call good
arXiv:2608.30466Self-authored retrieval can “lock” RAG into collapsed answers—at scale.
arXiv:2608.22118AI summaries shift clicks away from publishers and toward competitor engines—at the cost of user trust
arXiv:2608.18352How extraction siphons the web until it can’t renew itself
arXiv:2608.15896Most venues are missing from AI answers—what a true census audit reveals
arXiv:2608.07069Where grounded LLMs actually start naming people—and why most “visibility” measurements miss
arXiv:2607.23893One misleading “direct claim” can derail an agent that’s otherwise correct
arXiv:2607.17291Why citation coverage is sparse—and what actually predicts which sources get surfaced
arXiv:2607.15771Why GEO gains don’t translate into durable visibility (and which levers actually hold).
arXiv:2607.14035Variance components explain why brand answers won’t stabilize—and what to sample instead
arXiv:2607.13304When open-web search boosts coverage, it quietly degrades source trust
arXiv:2607.05217What LLMs treat as “brand reputation” is mostly other people’s web pages, not the brands themselves
arXiv:2606.25787Who gets the “top pick” in AI recommendations—and how consistent is it across models?
arXiv:2606.23057English-language prompts create a measurable “local-visibility” blind spot across languages
arXiv:2606.23165Why “mention” counts fail: fabricated citations scale differently by entity and query context
arXiv:2606.21595Why the first-run visibility gap between big brands and everyone else is so persistent
arXiv:2606.20065How LLM recommenders lock in incumbent brands—and the small tweaks that break it
arXiv:2606.17443When “search-and-answer” becomes endorsement for sale
arXiv:2606.16821Why one poisoned search result is enough to hijack LLM recommendations
arXiv:2606.13610When an assistant says the brand name: what actually moves browsing
arXiv:2606.10907Why safety alignment can turn retrieval-time injections into brand-level anti-promoters
arXiv:2606.09204FullCite turns inline citation into a document-plus-span problem, sharply improving quote-level grounding on ASQA.
arXiv:2606.07130Why “2x on ChatGPT” stories can be misleading without a tailwind control
arXiv:2606.04362When Cultural Knowledge Doesn't Transfer to Cultural Reasoning
arXiv:2606.01879How Prompt Language Rewrites Cultural Knowledge Before the Model Even Answers
arXiv:2605.30481Where LLM Fact-Checkers Go Wrong on Sources
arXiv:2605.30241When the Agent Learns From Its Own Corrections
arXiv:2606.02215Google AI Overviews drove a 12% rise in Reddit engagement, but AI Mode reversed those gains for experiential communities by substituting conversation for human discussion.
arXiv:2605.16428Coordinating a multi-page evidence ecosystem raises LLM search agent recommendation rates by up to 31 percentage points over the best single-page GEO baseline.
arXiv:2605.12887An audit of ChatGPT, Copilot, Gemini, and Perplexity finds ~16% of cited sources are AI-generated — with Copilot citing synthetic content in nearly 3 of every 10 citations.
arXiv:2605.23684Six frontier LLMs hallucinate 12–38% of scientific citations; a new agentic retrieval system hits zero hallucination at 30% better F1 and $0.05 per query.
arXiv:2605.14306Cosmetic prompt rewording drops AI brand-recommendation overlap by 21–32 percentage points — more divergence than switching providers entirely, across 12,000 runs.
arXiv:2605.27440Baidu's Aurora-Expiry uses RAG-augmented LLMs to infer query-specific expiration thresholds, cutting median document age 12.81% for time-sensitive queries in a 14-day live A/B test.
arXiv:2605.13052A 37,000-run audit of 533 brands finds RAG preserves the brand hierarchy: L4–L5 specialists face 48–52% invisibility while L1 leaders surface universally but convert at only 25–41%.
arXiv:2605.27439Schema.org markup gives retrieval agents 65.7% higher FAIR-compliant precision — but cuts query coverage by 29% where publishers haven't adopted it.
arXiv:2605.28787A Microsoft study shows a single false top search result drops GPT-5 accuracy from 65% to 18% — while humans solve the same queries at 93% — exposing a critical gap in agentic RAG deployments.
arXiv:2603.00801A new SoK survey of 118 works shows that agentic RAG's iterative retrieval and memory systems introduce failure modes that static metrics and current benchmarks cannot detect.
arXiv:2603.07379A new paper from Virginia Tech maps four failure modes that prevent pages from being cited in AI-generated responses. 43% of relevant pages receive zero citations under baseline conditions.
arXiv:2603.09296No LLM verifies even half its citations under any tested condition — and adding temporal cutoffs or other deployment constraints collapses verifiability to near zero.
arXiv:2603.07287A new statistical framework shows that single-run citation share metrics from Perplexity, SearchGPT, and Gemini carry confidence intervals wide enough to make most apparent SEO gains statistically indistinguishable from noise.
arXiv:2603.08924Rewriting pages as entity documents lifted AI answer accuracy ~30%; adding JSON-LD did almost nothing — it gets cut before indexing.
arXiv:2603.10700