The practical brief

The finding

Cited AI answers reduced measured article consumption per search by 18.2%, while article consumption per assigned reader increased by 5.9%. [c1] [c5] [c8]

Why it matters

A search-level decline can coexist with higher overall engagement; the denominator changes the business interpretation.

Search-level losses differed from reader-level outcomes
OutcomeAI-interface change versus Control
Search frequency+29.4% searches per assigned reader; 95% CI: +26.5% to +32.6% [c8]
Article consumption per search−18.2% topic-measured article consumption per search; 95% CI: −21.5% to −14.9% [c8]
Article consumption per reader+5.9% topic-measured article consumption per assigned reader; 95% CI: +1.7% to +10.1% [c8]
Relative Treatment-versus-Control changes within the Washington Post experiment, with reported 95% confidence intervals. Article consumption is topic-measured exposure, not verified reading. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: evaluate per-search and per-reader outcomes together, keeping answer exposure separate from article readership and business outcomes.

Study boundary

This single-publisher working paper does not establish external AI referrals, conversion gains or improved credibility.

Judge the reader journey, not just each search

If you operate a searchable content site, adding AI answers could reduce article consumption on each search without reducing it across readers’ activity. The business risk is judging a rollout by one denominator—or treating answer exposure as equivalent to readership.

This experiment found 29.4% more searches per assigned reader, 18.2% less measured article consumption per search and 5.9% more article consumption per assigned reader. Including answers, total information consumption rose 40.8% per search and 82.2% per reader. These are relative changes in group-total ratios, not percentage-point changes or proof that extra searching caused the gain.

Evidence [c4] [c8]

What the experiment changed and counted

The October 1–November 7, 2024 experiment included 37,561 cookie-based readers: 26,639 Control and 10,922 Treatment. Both groups searched the same Washington Post archive using the same retrieval infrastructure. Treatment added an AI answer with up to five citations above conventional results.

Answer content, source selection, citations and placement were tested together, not independently. “Consumption” means information provided through article opens and answer displays—not confirmed reading. Each measured item’s count was split across modeled news topics; answer topics were primarily inferred from cited articles.

Evidence [c1] [c5] [c11]

Topic exposure broadened, including beyond answer displays

Including answers, the individual “effective topic count”—a diversity measure reflecting how evenly exposure is distributed—rose from 4.58 to 5.55. Across the audience, it rose from 44.37 to 46.42. Total article-and-answer consumption shifted toward less-popular topics at both levels.

Opened articles also shifted toward less-popular topics at both individual and audience-wide levels, and audience-wide article consumption became less concentrated. However, individual article-only concentration differences were not statistically significant. The individual analyses required at least two opens or displays, potentially selecting different readers after assignment rather than comparing everyone assigned.

Study-defined shared topics reached 55.4% of Treatment readers versus 36.9% of Control readers when answers counted. Treatment’s article-only reach was 35.1%: the increased reach came from including answers. Shared topics do not necessarily mean shared facts or viewpoints.

Conventional-result article opens were the next action after 52.9% of Control searches versus 32.1% of Treatment searches; cited-source opens accounted for 14.3% in Treatment. Session-ending shares stayed near 21%. These are immediate-next-action shares, not total readership.

Evidence [c2] [c3] [c7] [c9] [c14] [c18] [c20] [c25]

A realistic interpretation for an on-site pilot

Our interpretation is to assess the whole journey. Lower article consumption per search could miss repeat searching; higher answer-inclusive exposure could overstate readership. The findings support a measurement distinction, not a proven growth tactic.

  • Track per-search and per-reader engagement separately.
  • Keep answer displays distinct from article opens.
  • Measure conversion outcomes separately rather than inferring them from topic exposure.

Evidence [c1] [c5] [c8]

What this study cannot tell your business

This archive-only experiment cannot establish whether external assistants will cite your business, send visitors or improve credibility. Broader news-topic consumption does not demonstrate viewpoint diversity, factual accuracy or trust.

The working paper’s causal interpretation depends on same-day groups being comparable after accounting for measured pre-search characteristics. Assignment records were unavailable, and individual diversity analyses select readers by activity after assignment. Robustness checks support the broad findings, but effect sizes depend on how answer exposure is represented and weighted.

Evidence [c1] [c12] [c14] [c16] [c20] [c21]

The source and evidence

Breaking News Out of the Filter Bubble: Generative AI Search Diversifies Collective Attention and Raises Shared Information Consumption

Heeseung Andrew Lee, Dokyun Lee, Gwanhoo Lee, Dongwon Lee · 2026-09-30

Study limitations and disclosures

  • This 38-day, single-publisher working paper tests a bundled interface, not separate answer or citation effects. It does not establish external referrals or conversion outcomes.
  • Assignment code, allocation logs and the intended ratio were unavailable. Causal interpretation assumes same-day groups are otherwise comparable after accounting for measured pre-search characteristics. Differences were small but jointly statistically significant; adjusted estimates were close to original differences. Cookie-based readers are not necessarily unique people.
  • Opens and displays do not demonstrate reading or comprehension. Unmapped content is excluded from topic consumption. Individual diversity analyses condition on post-assignment eligibility; robustness checks preserve broad findings, but effect sizes depend on answer-topic representation and weighting.
  • The shared-information/engagement and concentration/popularity comparison families were not preregistered.
  • Author affiliations include the University of Texas at Dallas, Boston University, American University and Hong Kong University of Science and Technology. The Washington Post operated the experiment and supplied restricted de-identified data; authors performed secondary analysis. They report independent interpretation, publisher review limited to proprietary information, no publisher employment or financial interests, no other competing interests and no external funding.
  • The authors disclose using Claude Fable 5 and GPT-6 via Codex for topic-model development and review, replication analyses and text editing. They state that they verified all analyses.

Evidence behind this briefing

[c1] The publisher’s randomized field experiment assigned readers by browser cookie at their first on-site search during October 1–November 7, 2024. The analytic sample comprised 26,639 Control and 10,922 Treatment readers. Treatment added an AI answer with up to five citations above conventional results, using the same corpus and retrieval infrastructure.

Materials and Methods — Setting and design · Read the study

[c2] Shared-topic information provided per assigned reader increased from 0.676 to 1.144, approximately 69%, while non-core information increased from 0.687 to 1.339; both comparisons had Holm-adjusted P < 0.001. These are summed topic-share units, not verified comprehension. Shared-topic reach increased from 36.9% to 55.4% when answers were included, whereas article-only reach decreased from 36.9% to 35.1%. Answers accounted for 98.0% of the shared-information increase, a descriptive decomposition rather than an isolated causal effect of answer content.

Supplementary material — S3; supporting quantities and qualifications in S2/Table S1 and S5/Table S5 · Read the study

[c3] Total information became less concentrated and shifted toward less-popular topics at individual and aggregate levels. Individual effective topic counts increased from 4.58 to 5.55, and aggregate counts from 44.37 to 46.42; differences were +0.97 (95% CI 0.86–1.08) and +2.05 (1.46–2.58). Individual Top-10 share fell 4.0 percentage points (95% CI −5.0 to −3.0), and aggregate share fell 4.4 points (−5.4 to −3.4); all eight Table 1 comparisons had Holm-adjusted P < 0.001. Individual estimates require at least two article opens or answer displays. Article-only individual concentration did not differ significantly.

Results — AI answers and cited articles change how readers consume; Table 1 and its notes · Read the study

[c4] Treatment changed the next action after search: conventional-result clicks and browsing fell by 21 and 5 percentage points, while cited-source clicks and follow-up searches without an article click rose by 14 and 11 points. Approximately 21% of searches ended sessions in each group. Despite 18% lower article consumption per search, 29% more searches yielded approximately 6% higher article consumption per reader. These outcomes concern the bundled interface and on-site activity, not publisher referrals from third-party AI services.

Results — AI answers and cited articles change how readers consume · Read the study

[c5] Consumption is an exposure proxy: opened articles and displayed answers count without establishing reading or understanding. Answer topics are inferred by equally averaging cited articles’ topic distributions rather than directly analyzing answer text. Unmapped opens and displays remain in engagement records but are excluded from topic consumption.

Materials and Methods — Measuring consumption · Read the study

[c6] The Washington Post designed and operated the experiment and supplied restricted de-identified records. The authors state that interpretation is their own, publisher review is limited to proprietary information, and they have no employment or financial interests in the publisher or other competing interests.

Competing interests · Read the study

[c7] Across 46,508 Control and 24,683 Treatment searches, the next action shifted from conventional-result articles toward cited-source articles and follow-up searches. Conventional-result opens fell 20.8 percentage points (95% CI −21.82 to −19.75); cited-source opens rose 14.3 points (13.73 to 14.86). Session-ending shares were essentially unchanged. These are immediate-next-action shares, not total article readership.

S17, Table S24 · Read the study

[c8] Treatment had 29.4% more searches per assigned reader (95% CI 26.5–32.6), 18.2% less measured article consumption per search (−21.5 to −14.9), but 5.9% more measured article consumption per reader (1.7–10.1). Including AI answers increased measured total consumption per reader by 82.2% (76.7–88.0). These group-total ratios do not isolate the causal contribution of additional searches; consumption excludes content without topic estimates.

S18, opening paragraph and Table S25 · Read the study

[c9] Counting only opened articles, Treatment shifted toward less-popular topics at individual and aggregate levels, but individual concentration differences were not statistically significant. Individual Gini difference was −0.0003 (95% CI −0.0021 to 0.0014; Holm P = 1.000), while individual top-10-topic share fell 2.58 percentage points (−3.86 to −1.31; Holm P = 0.002). Individual estimates use 7,992 Control and 3,423 Treatment readers with at least two opens and some topic consumption.

S13, opening paragraph and Table S17 · Read the study

[c10] The less-popular-topic consumption shift was not mirrored by requested topics or same-query offered lists. Neither cited-source versus conventional lists nor combined versus conventional lists showed statistically significant popularity differences. These paired comparisons included 3,253 and 4,114 queries respectively, with incomplete and unequal search-list coverage; lack of significance does not establish equivalence.

S16, “Articles offered for the same query,” Table S23 · Read the study

[c11] Topic-based consumption gives each measured article open or substantive answer display one count split across topics, starting at the first search. Main answer-topic estimates use equal-weight averages of cited articles, not direct measurement of answer facts or viewpoints; four excluded topics are removed and remaining shares normalized. Alternative answer-text, citation-order and half/double-answer-weight checks preserve the broad findings, although effect sizes depend on representation.

S24, answer-topic measurement; S12, Tables S15–S16 · Read the study

[c12] Assignment implementation was not independently disclosed: the publisher reports cookie assignment at first search, but assignment code, allocation logs and intended allocation ratio were unavailable. Treatment shares varied across entry days (P < 0.001), so main permutation tests preserve group sizes within each of 38 first-search days. The final sample contains 26,639 Control and 10,922 Treatment cookie-based readers; interpretation depends on assignment assumptions referenced outside this chunk.

S19, opening paragraph, Table S26; S30 · Read the study

[c13] The table measures reach using 18,571 Control article readers. Consumption shares total 100% before rounding, reach percentages may overlap, and topic IDs preserve the model’s original numbering.

Topic table, Note immediately before References · Read the study

[c14] The study examines consumption diversity across all news topics, rather than only political ideology; the footnote explains that ideological measures apply only to political content.

Footnotes, Footnote 1 · Read the study

[c15] The title page explicitly establishes author institutional affiliations: University of Texas at Dallas, Boston University (two affiliations), American University, and Hong Kong University of Science and Technology. The statement that institutional affiliations are not established should be removed.

Title page — numbered affiliations 1–5 · Read the study

[c16] Causal interpretation assumes that same-day Treatment and Control readers are otherwise comparable after accounting for measured pre-search characteristics. Those characteristics differed only slightly, but the differences were jointly statistically significant; adjustment for characteristics and first-search week yielded estimates close to the original group differences.

Materials and Methods — Setting and design, second paragraph · Read the study

[c17] The authors disclose using Claude Fable 5 and GPT-6 via Codex to assist with developing and reviewing BERTopic topic modeling and replication analyses, and with text editing. They state that they verified all analyses.

Materials and Methods — Inference, final paragraph · Read the study

[c18] The main individual-level analysis is conditional on consuming at least two articles or AI answers, not an unconditional comparison of all assigned readers. Because counting answers lets more Treatment readers meet this post-assignment requirement, eligibility can select different readers across groups.

Supplementary material — S9, second paragraph · Read the study

[c19] Answer-topic consumption is represented by equally averaging the topic distributions of cited articles, excluding four topics, and rescaling the remaining shares. This is a representation of information provided, not a direct measurement of how fully readers read or understood the answer.

Materials and Methods — Measuring consumption, first paragraph · Read the study

[c20] The individual diversity checks do not eliminate differences between readers selected by consumption after assignment; these results should not be presented as an unconditional comparison of all assigned readers.

S9, paragraph immediately before S10 · Read the study

[c21] The broad shared-information, concentration and popularity findings persist under alternative answer measurements and weights, but effect sizes depend on how answers are measured. Checks include answer-text topics, citation-order weighting, matched answer coverage, and counting answers as half or twice an article.

S12, opening two paragraphs; Tables S15–S16 · Read the study

[c22] Pre-search characteristics explained little assignment variation (R2 = 0.005), but their joint association with assignment was statistically significant (P < 0.001). Estimates remained similar after adjustment for these characteristics and entry week; this does not itself establish the causal comparability assumption.

S20, paragraph following Table S28 · Read the study

[c23] The 11 shared-information and engagement comparisons were selected before multiplicity adjustment but were not preregistered. One conditional cosine-similarity comparison was not tabulated because its direction depended on topic representation, although it remained in the adjustment.

S29, first paragraph · Read the study

[c24] The eight concentration and popularity comparisons in Table 1 were not preregistered. Their tests used 100,000 permutations and Holm adjustment; bootstrap confidence intervals held the observed Control popularity ranking fixed, whereas permutations rebuilt it.

S31, first paragraph · Read the study

[c25] When only opened articles are counted, consumption shifts toward less-popular topics at both levels and aggregate concentration decreases, but individual concentration differences are small and not statistically significant. Thus individual diversification of total article-and-answer consumption should not be generalized to article consumption alone.

S13, opening paragraph; Table S17 · Read the study