The practical brief
The finding
For one contested historical event, just 2 of 119 unique citations appeared in more than one of three Wikipedia language editions; explicit disagreement-preserving prompts also improved model-judged summary balance in a separate test. [c1] [c5] [c6] [c7] [c8]
Why it matters
For businesses reviewing multilingual AI content, checking wording alone may miss differences in the underlying sources.
| Prompt strategy | LLM-judged neutrality | LLM-judged conflict preservation |
|---|---|---|
| Standard | 0.47 on a 0–1 scale; Night Attack summary test | 0.60 on a 0–1 scale; Night Attack summary test [c5] [c6] |
| Mediator | 0.95 on a 0–1 scale; Night Attack summary test | 1.00 on a 0–1 scale; Night Attack summary test [c5] [c6] |
| Academic | 0.97 on a 0–1 scale; Night Attack summary test | 1.00 on a 0–1 scale; Night Attack summary test [c5] [c6] |
| Disagreement-preserving (“adversarial”) | 1.00 on a 0–1 scale; Night Attack summary test | 1.00 on a 0–1 scale; Night Attack summary test [c5] [c6] |
What to try
BLURSOR’s practical interpretation
Our interpretation: compare source lists across relevant languages and ask AI-assisted drafts to attribute unresolved disagreements, followed by factual review.
Study boundary
This historical case study does not establish effects on business discoverability, customer trust, citations or sales.
The risk: consistent wording can hide inconsistent sources
Before approving multilingual AI-assisted content, decide whether you need a source-level review as well as a translation check. This study illustrates why: accounts of the same event can draw on largely separate evidence pools, even when their language sounds neutral.
That is a credibility concern, not a demonstrated marketing outcome. The researchers studied contested history on Wikipedia, not brand recommendations or commercial search. Its practical relevance is to how businesses inspect sources and represent disagreement—not how they win AI visibility.
What the researchers actually compared
The study covered three events in Romanian history across Romanian, Hungarian, Russian, Turkish and English Wikipedia editions. Its citation analysis examined the Battle of Posada in Romanian, Hungarian and English. A separate temporal analysis covered 29 article versions from 2005–2024.
For Posada, Gemini 3.0 Pro Preview classified citations by narrative stance: Pro-Romanian, Pro-Hungarian, Neutral or Contested. Two human annotators checked a random half of Romanian and English citations and all Hungarian citations. Human–model agreement was substantial, but agreement does not establish historical truth.
Only 2 of 119 unique citations were shared across more than one edition. The researchers call this “citation isolation”: different language communities relying on different sources. They classified 91.1% of Romanian-edition citations as Pro-Romanian, compared with 23.1% in Hungarian and 31.8% in English. These percentages describe narrative stance, not factual accuracy; the prose does not report each edition’s citation denominator.
Preserving disagreement improved summary scores, not verified truth
A separate experiment extracted conflicting Romanian and Turkish claims about the Night Attack at Târgoviște and compared four summary prompts. Another AI instance scored the summaries for neutrality, conflict preservation and source attribution.
Here, neutrality meant perspectival balance: representing and attributing disagreement without endorsing one side. The “adversarial” prompt was not an attack; it explicitly instructed the model to preserve conflict rather than resolve it. It received higher scores than standard summarization, although academic and mediator prompts also performed well.
This is evidence about model-assessed representation in one test. A maximum score does not mean perfect accuracy, and the paper does not report trial counts or uncertainty estimates.
A realistic interpretation for your content workflow
Our interpretation is to use multilingual comparison as a diagnostic, not a proven optimization tactic. Where claims are sensitive or contested, inspecting the evidence behind each language version may be more useful than checking whether translations match.
- Compare the underlying references, not just the translated prose. Different source lists merit inspection; they do not automatically prove an error.
- Ask AI-assisted drafts to identify which source makes each disputed claim rather than manufacturing a single consensus.
- Separate factual verification from fair representation. Attributing opposing claims does not make both valid.
What this study cannot tell your business
The findings cannot predict whether multilingual monitoring will increase recommendations, citations, revenue or trust. Coverage is narrow, and without an uncontested-topic comparison, the study cannot establish whether citation isolation is distinctive to contested history.
Both AI classification and human validation may reproduce national perspectives. The proposed geopolitical explanation for a Russian article’s 2024 shift is explicitly speculative. Treat the study as a reason to examine evidence and attribution—not as a verdict on which language edition is correct.
The source and evidence
The Rashomon Wikipedia: A Data-Perspectivist Analysis of Divergent Historical Narratives
Study limitations and disclosures
- The study covers three contested Eastern European historical events, not representative Wikipedia topics or commercial queries. It does not measure business visibility, customer trust or marketing performance, and lacks an uncontested-topic baseline.
- LLM stance classifications may contain bias. Native-speaker human annotators may share the national perspectives under investigation, so their agreement with the model is not an independent benchmark of historical truth. Edition-specific citation denominators are not supplied in the reported stance-percentage prose.
- Summary scores came from a separate LLM instance, not human evaluation or factual verification. The paper does not identify the summary-generation or judging model, report repeated-run counts or uncertainty estimates, or describe human validation of these scores.
- The proposed geopolitical explanation for the Russian Bessarabia article’s 2024 shift is speculative. The temporal method does not provide a complete numerical mapping from categorical stance confidence to its 0–100 narrative scores.
- The metrics assess perspectival balance, not historical truth. Preserving disagreement still requires fact-checking to avoid giving unsupported claims false equivalence. The study used public CC BY-SA 4.0 data and anonymized its human annotators.
- The authors list University of Bucharest, Romania affiliations, including the HLT Research Center, Interdisciplinary School of Doctoral Studies, Faculty of Mathematics and Computer Science, and Faculty of Foreign Languages and Literatures. Funding is acknowledged from the Romanian Hub for Artificial Intelligence project, MySMIS 351416, and the Ministry of Research, Innovation and Digitization/CNCS–UEFISCDI SIROLA grant, PN-IV-P1-PCE-2023-1701. The supplied paper contains no explicit commercial-interest declaration.
Evidence behind this briefing
[c1] For the Battle of Posada, only 2 of 119 unique citations were shared across more than one of the Romanian, Hungarian and English Wikipedia editions. This is evidence of source separation in this case, not a general measurement of Wikipedia or AI citation behavior.
Section 4.1, Experiment 1: Citation Isolation (Posada) · Read the study
[c2] In the 119-citation Posada analysis, 91.1% of Romanian-edition citations were classified as Pro-Romanian, versus 23.1% in Hungarian and 31.8% in English. These are classified narrative stances, not factual-accuracy scores; the prose does not supply each edition's citation denominator.
Section 4.1, Bias Scores · Read the study
[c3] Citation stance was classified using Gemini 3.0 Pro Preview with enhanced article context. Two human annotators validated a random 50% sample of Romanian and English citations and all Hungarian citations; human–LLM Cohen's kappa was 0.80, 0.75 and 0.85, respectively. Agreement is not an independent measure of historical truth.
Section 3.2.2, Human validation; model specified in Section 3.2.1 · Read the study
[c4] Across a temporal dataset of 29 article versions from 2005–2024, the Russian Bessarabia 'Accession' article received a Pro-Romanian narrative score of 72 in 2024 on the study's 0–100 scale. The suggested connection to post-2022 geopolitics is explicitly speculative, not an established causal result.
Section 4.2, Experiment 2: Temporal Evolution, Bessarabia (1940) · Read the study
[c5] The Peace-Maker pipeline extracted and matched conflicting claims, then compared four summary prompts. A separate LLM instance judged neutrality, conflict preservation and source attribution on 0–1 scales; these are model-assessed representational qualities, not human preferences or verified accuracy.
Section 3.3, The "Peace-Maker" Pipeline, step 4 · Read the study
[c6] For summaries of conflicting Romanian and Turkish Night Attack claims, the adversarial prompt scored 1.00 for both LLM-judged neutrality and conflict preservation, versus 0.47 and 0.60 for standard prompting. Academic and mediator prompts also scored highly. This supports preserving attributed disagreement in this test, not a claim of perfect factual accuracy; the supplied paper does not report trial counts or uncertainty estimates.
Section 4.3, Experiment 3: Peace-Maker LLM, Table 1 · Read the study
[c7] The authors disclose potential LLM classification bias, narrow geographic and event coverage, and possible circularity in human validation because native-speaker annotators may share the national perspectives being measured.
Limitations · Read the study
[c8] Preserving conflicting claims is not sufficient to establish their validity. The authors identify a need for fact-checking to prevent false balance, an important qualification for businesses seeking credible AI answers.
Section 5.3, AI as a Mediator, Not a Judge · Read the study