The practical brief

The finding

Across two scientific-answer datasets, human preferences were positively associated with citation diversity and negatively associated with total citation count; AI judges showed less consistent patterns. [c1] [c2] [c3] [c4] [c5] [c6] [c8] [c13] [c15]

Why it matters

A favorable AI evaluation is not necessarily a check that cited sources support an answer.

Human citation preferences in two scientific-answer datasets
Citation featureSciArena human preferencesResearchQA human preferences
Citation diversity+1.8 percentage points in selection probability per 1-SD predictor-difference increase; 12,249 comparisons+11.7 percentage points in selection probability per 1-SD predictor-difference increase; 169 comparisons [c1] [c2] [c4]
Total citation count−3.5 percentage points in selection probability per 1-SD predictor-difference increase; 12,249 comparisons−6.5 percentage points in selection probability per 1-SD predictor-difference increase; 169 comparisons [c1] [c2] [c4]
Measured observational associations per one-standard-deviation increase in the response-pair predictor difference—not effects of adding or removing one citation. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: review source support separately from answer preference, rather than treating citation volume or an AI judge’s approval as sufficient evidence.

Study boundary

These observational results depend on datasets, modeling assumptions and judge prompts. They do not test business discovery, customer trust or tactics for winning citations.

Separate a preferred answer from a verified answer

If your business uses AI to produce research summaries or score answers, do not mistake a citation-rich presentation for verified support. This study examines which scientific answers people and AI judges prefer—not whether those answers are accurate.

Researchers analyzed existing human comparisons and asked four open-weight language models—Llama, Gemma, Qwen and Skywork—to judge the same answer pairs. The shared prompt assigned an expert-in-scientific-literature-synthesis role. Asking judges to consider accuracy and citation use was not a separate test of factual accuracy.

Evidence [c1] [c2] [c4] [c14] [c17]

Two datasets, with very different breadth

SciArena supplied 12,249 eligible comparisons of answers from 23 models, using majority preferences from 102 researchers. Both answers needed citations; ties and “both bad” judgments were excluded. ResearchQA supplied 169 clear-preference examples comparing two answer-generating models, with annotators drawn from 31 PhDs.

The researchers used mixed-effects models: statistical models that account for group differences, such as question subjects and answer-generating models. They controlled for answer length and semantic-content differences and tested both answer orders.

Those adjustments involve choices. Random-effects selection lacks consensus; the researchers started with a maximal feasible structure, removed terms to obtain a non-singular fit, and simplified using AIC, a model-comparison criterion. They acknowledge possible over-specification.

Evidence [c1] [c2] [c4] [c8] [c9]

Citation variety and volume pulled apart

On SciArena, greater citation diversity was associated with a 1.8-percentage-point increase in human selection probability, while greater total citation count was associated with a 3.5-point decrease. These estimates describe a one-standard-deviation increase in the difference between paired answers—not adding one citation.

All four AI judges showed larger positive diversity associations on SciArena, ranging from 5.2 to 7.5 percentage points on the same scale. Citation-count patterns varied. On ResearchQA, significant citation-feature associations appeared only for diversity in Llama and Gemma.

ResearchQA humans also penalized citation mismatch: disagreement between inline citations and reference-list entries. That measures consistency, not independently verified fabrication. AI judges did not reproduce humans’ strongly negative mismatch association.

Evidence [c1] [c2] [c3] [c4] [c5]

Use preference scores for a narrower purpose

Our interpretation is to keep answer preference and evidence quality as separate review questions. Favoring citation patterns does not establish whether a source exists, supports the claim or applies to your situation.

The study does not justify trimming references to an ideal count or adding varied sources to improve business visibility. Neither intervention was tested.

  • Review readable presentation separately from verified source support.
  • Check inline citations against reference lists, then inspect the underlying sources.
  • Treat an AI preference score as an answer judgment, not a substitute for source verification.

Evidence [c3] [c4] [c5] [c6] [c17]

What the study cannot tell your business

The results are correlations, not causal effects. They do not measure customer trust, sales, business recommendations or website citation probability. ResearchQA’s small sample and single model pair limit transfer to commercial answers; cross-dataset differences remain unexplained.

Prompt choice also matters. All AI judges used a shared scientific-literature-synthesis prompt; substantially different prompts or personas could produce different preferences. Additional instructions against length and order bias were tested only with Qwen3-30B. Their high correlation with its original results does not establish robustness across all judges or substantially different prompts.

The statistical models assume linear relationships. Although the authors found no evidence of violation, nonlinear relationships and interactions remain possible. This and random-effects selection qualify the reported associations.

The authors disclose AI assistance with writing refinement, figure brainstorming and coding, while accepting responsibility. Code and data release was promised upon publication, not established as already available.

Evidence [c6] [c8] [c9] [c11] [c12] [c13] [c14] [c15]

The source and evidence

On the Role of Citations in Preference Data

Yu Hou, Hal Daumé III, Rachel Rudinger, William Walden · 2026-07-06

Study limitations and disclosures

  • The analysis is observational: citation features correlate with preferences, but changing those features was not shown to cause preference changes.
  • The evidence concerns scientific-answer comparisons, not commercial discovery or customer behavior. ResearchQA contains only 169 eligible comparisons between two answer-generating models; cross-dataset differences in human preferences remain unexplained.
  • AI-judge preference does not establish citation accuracy. Citation mismatch measures disagreement between inline citations and reference-list entries, not independently confirmed hallucination.
  • All AI judges used a shared scientific-literature-synthesis prompt. Different prompts or personas could change their preferences. Additional length- and order-bias instructions were tested only with Qwen3-30B; high correlation with its original results does not establish robustness across all judges or substantially different prompts.
  • Random-effects selection lacks consensus. The maximal-first structure was simplified to obtain a non-singular fit and using AIC; the authors acknowledge possible over-specification. Reported associations depend on these modeling choices.
  • Fixed-effect predictors assume linear relationships. Although the authors found no evidence of violation, nonlinear relationships and interaction effects remain possible and were left for future investigation.
  • The authors disclose AI assistance with writing refinement, figure brainstorming and coding, while accepting responsibility for the work. Code and data release was promised upon publication rather than established as currently available.

Evidence behind this briefing

[c1] The study elicited pairwise judgments from Llama3.3-70B, Gemma3-27B, Qwen3-30B, and Skywork-Critic-70B, testing both response orders: 24,498 judgments per model on SciArena and 338 on ResearchQA. Mixed-effects logistic models controlled for response length and semantic-content differences; reported effects are percentage-point changes in selection probability per one-standard-deviation increase in a standardized predictor difference.

Section 2, LLM Inference Experiment; Fixed-Effects Predictors; Average Marginal Effects · Read the study

[c2] On SciArena, higher citation diversity was associated with a 1.8-percentage-point increase in human selection probability, while higher total citation count was associated with a 3.5-point decrease, each per one-standard-deviation increase in the response-pair difference. The filtered sample comprised 12,249 comparisons across 23 response-generating models and 480 unique model pairs, using majority preferences from 102 researchers; both responses needed citations, and ties and both-bad labels were excluded. Older-citation preference was not significant.

Section 3.2, Human Preferences; sample qualification in Section 3.1, Data · Read the study

[c3] On the same SciArena comparisons, the four LLM judges had larger positive citation-diversity associations than humans: 7.5, 5.2, 6.2, and 7.0 percentage points for Llama, Gemma, Qwen, and Skywork, respectively. Citation-count results varied: Llama −2.1 points and Skywork −12.8 points (p<0.001), with nonsignificant results for Gemma and Qwen. No judge showed a significant citation-age preference. These are preference associations, not evidence that the citations were accurate or verified.

Section 3.3, LLM Judge Preferences · Read the study

[c4] ResearchQA human preferences showed positive citation-diversity and negative citation-count and citation-mismatch associations: +11.7, −6.5, and −10.1 percentage points per one-standard-deviation predictor-difference increase. Mismatch was significantly negatively correlated with selection. Response length had a +21.4-point association; semantic-content difference was not significant. Scope was only 169 clear-preference examples, with annotators from a pool of 31 PhDs, comparing gpt-4.1-mini against gemini-2.5-flash across eight scientific fields in five domains. Mismatch means disagreement between inline citations and reference-list entries, not independently verified hallucination.

Section 4.2, Human Preferences; sample qualification in Section 4.1, Data · Read the study

[c5] On ResearchQA, significant citation-feature associations for LLM judges were confined to citation diversity for Llama (+7.8 percentage points) and Gemma (+8.8 points). Qwen and Skywork diversity estimates were positive but nonsignificant. The judges did not reproduce humans’ strongly negative citation-count and mismatch associations, showing that results depend on the dataset and judge rather than a universal citation preference.

Section 4.3, LLM Judge Preferences · Read the study

[c6] The authors explicitly characterize the fitted effects as correlations, not causal effects of citation features on human or LLM preferences.

Limitations, Correlation vs. Causality · Read the study

[c7] Zero-shot judges were Llama-3.3-70B-Instruct, Gemma-3-27b-it, Qwen3-30B-A3B-Instruct-2507 and Skywork-Critic-Llama-3.1-70B, using the SciArena prompt and temperature 0. Appendix A.2 additionally describes reversing response order to control potential positional bias in citation preferences.

Appendix A.1–A.2 · Read the study

[c8] SciArena mixed-effects modeling began with the maximal feasible random-effects structure, removed terms for a non-singular fit and simplified using AIC. The authors acknowledge no consensus on random-effects selection, computational cost and likely over-specification. Data groups included 23 models, six query types, 25 query subjects in five categories and 480 model pairs; these are group counts, not a response-observation denominator.

Appendix B, opening paragraphs · Read the study

[c9] ResearchQA had only 169 observations, responses from gpt-4.1-mini and gemini-2.5-flash, and eight fields. Its maximal random-effects structure was therefore restricted to a field-level random intercept, a narrower scope than the SciArena model.

Appendix B, ResearchQA model specification · Read the study

[c10] Citation dates were log-transformed because their distribution was heavily skewed: 79.67% of citations fell after 2019, with median year 2023 and modal year 2024. The footnote does not state the total citation count or explicitly identify the dataset for these figures. BERTScore used roberta-large.

Footnotes 4–5 · Read the study

[c11] Observed human preferences for response length and content differed from SciArena. The authors propose differences in annotator populations, collection setups and response comparisons as possible explanations, not established causes, and defer investigation.

Footnote 7 · Read the study

[c12] The authors disclose LLM assistance with writing refinement, figure brainstorming and coding, while denying generation of content or sentences from scratch and accepting responsibility. Code and data release is promised upon publication rather than established as already available.

Appendix C; Footnote 1 · Read the study

[c13] The authors assume linearity for the fixed-effects predictors. They found no evidence that this assumption was violated, but acknowledge that more complex nonlinear relationships may exist and leave potential nonlinear and interaction effects for future investigation. This qualifies interpretation of the reported per-standard-deviation preference associations.

Limitations — Linearity Assumption · Read the study

[c14] The LLM judges used the same prompt for SciArena and ResearchQA, adopting an expert-in-scientific-literature-synthesis persona and asking for preferences between citation-attributed responses.

Section 2 — General Setup: LLM Inference Experiment and LLM Judge Prompt · Read the study

[c15] The authors disclose that prompts and personas could change observed LLM preferences. Additional length- and order-bias instructions were tested only with Qwen3-30B; their results correlated highly with the original results, but this does not establish robustness across all judges or substantially different prompts.

Limitations — Prompting · Read the study

[c16] The four LLM judges used the same SciArena prompt in zero-shot experiments: Llama-3.3-70B-Instruct, Gemma-3-27b-it, Qwen3-30B-A3B-Instruct-2507 and Skywork-Critic-Llama-3.1-70B.

Appendix A.1 — LLM Judge Experiment Setup · Read the study

[c17] The shared judging template assigned a scientific-literature-synthesis expert role and asked judges to select their preferred response based on relevance, accuracy, clarity and citation use. This is a response-preference task, not a separate measurement of factual accuracy.

Appendix A.2 — LLM Judge Prompt (Flipped Order) · Read the study

[c18] The flipped-order experiment retained the same prompt template while reversing response presentation order; it addresses positional bias rather than sensitivity to substantially different prompts or personas.

Appendix A.2 — Flipped Response Order Experiment · Read the study