The practical brief

The finding

Intent guidance improved some scientific-report scores, but results depended on the model, task, and whether guidance covered paragraphs alone or citations too. [c1] [c2] [c4] [c6] [c10]

Why it matters

Better attribution in reports built from supplied evidence does not establish that more customers will discover your business in AI answers.

The type of guidance changed the outcome
Model and evaluationBaselineIntent-aware variant
Gemini 2.5 Pro: SQA-CS-V2 overall88.1 overall score points on SQA-CS-V289.7 overall score points on SQA-CS-V2 with intent prompting [c2]
qwen3-8b: SQA-CS-V2 overall83.2 overall score points on SQA-CS-V2 with baseline fine-tuning88.6 overall score points on SQA-CS-V2 with intent-multiview fine-tuning [c3]
o3: DeepScholar overall46.8 overall score points on modified DeepScholar43.2 overall score points on modified DeepScholar with combined paragraph-and-citation intents [c4] [c10]
o3: DeepScholar paragraph-only recovery46.8 overall score points on modified DeepScholar49.3 overall score points on modified DeepScholar with paragraph-only intents [c10]
Measured benchmark scores, not business outcomes. Different tasks are not directly comparable; the o3 rows distinguish combined guidance from paragraph-only guidance. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: compare paragraph-only guidance with combined paragraph-and-citation guidance, keeping sources fixed and checking claim support separately from readability.

Study boundary

The experiments tested scientific writing, not brand inclusion, commercial source discovery, conversions, or customer trust.

Judge report quality separately from business visibility

If you are choosing an AI tool for market research or evidence-backed content, do not treat stronger citation scores as proof that your business will appear more often in AI answers. This study supports a narrower decision: whether explicit guidance about the purpose of writing can improve reports built from supplied evidence.

The researchers used structured tags for paragraph and citation “intents”—the functional reason a passage or reference belongs in a report. They tested intent-aware prompting and trained smaller models on synthetic reports containing those intents. This changes the writing process; it is not a tested optimization for your website.

Evidence [c1] [c3] [c6]

The experiment isolated writing from source discovery

The evaluations covered scientific question answering and related-work writing. SQA-CS-V2 had 100 test samples, modified DeepScholar covered 63 papers, and ResearchQA’s AI subset had 50 test questions. Retrieved material was fixed for each query to isolate writing performance. ResearchQA used no retrieval and only paragraph intents.

On SQA-CS-V2, Gemini’s overall score rose from 88.1 to 89.7 points. Citation precision—whether references support their associated claims—and citation recall—whether claims needing support receive it—improved, while its rubric score stayed unchanged. Better attribution therefore did not mean broader coverage of the requested answer.

For qwen3-8b, intent-multiview supervised fine-tuning, meaning training on multiple instruction-and-answer versions of each report, reached 88.6 overall versus 83.2 for baseline fine-tuning. Intent-trained models also produced intents during use, so the comparison was not an isolated training-data change.

Evidence [c1] [c2] [c3]

Paragraph guidance and citation guidance behaved differently

On DeepScholar, o3’s overall score fell from 46.8 to 43.2 with combined paragraph-and-citation intents, and citation precision fell from 39.1 to 27.2. But paragraph-only guidance recovered the overall score to 49.3, kept citation precision at 39.1, and raised claim coverage from 40.2 to 44.3. The practical lesson is not simply that intent guidance helps or hurts: which annotations you use matters.

A qualitative inspection of just 20 o3 claims found added information from model memory in nearly 60% of cases. Explaining why a citation belongs cannot substitute for checking what its source says.

A separate study of 20 participants and 71 reports found higher self-reported usefulness for deciding which paragraphs or citations to inspect. Recruitment targeted computer-science postgraduate expertise. These ratings measured reading guidance and confidence, not verified comprehension or factual accuracy.

  • Our interpretation: pilot paragraph-only annotations separately from combined paragraph-and-citation annotations before adopting either.
  • Keep the source packet unchanged and assess claim support separately from readability.
  • Treat website-content applications as untested extensions, not proven routes to AI recommendations.

Evidence [c4] [c5] [c10] [c11] [c12]

What the study cannot tell your business

The study does not measure brand visibility, source ranking, conversions, or customer trust. Its mainly English scientific tasks may not transfer to commercial content. The intent schema was synthetic rather than grounded in observations of human writing; flexible schemas also bring a consistency trade-off.

Benchmark scoring used AI judges. The authors used a significance threshold of 0.1; o3’s reported SQA-CS-V2 improvement did not meet the stricter 0.05 threshold. Neither benchmark gains nor reader preferences establish reliable business deployment.

Disclosures include ONR award and Amazon AI PhD Fellowship support for Xinran Zhao, plus LLM assistance for grammar correction and coding—not paper writing or overall code-base logic. The authors report open-sourced code, training data, and checkpoints; release contents were not independently verified here. Their low-risk assessment concerns human-verified report queries, not general business use.

Evidence [c2] [c6] [c7] [c8] [c9]

The source and evidence

Improving Attributed Long-form Question Answering with Intent Awareness

Xinran Zhao, Aakanksha Naik, Jay DeYoung, Joseph Chee Chang, Jena D. Hwang, Tongshuang Wu, Varsha Kishore · 2026-03-28

Study limitations and disclosures

  • No brand visibility, marketing conversion, source-ranking, or commercial discoverability outcomes are tested.
  • Fixed retrieval prevents attributing these results to improved source discovery; ResearchQA uses no retrieval.
  • The body’s broad improvement language is qualified by the o3 DeepScholar regression and by gains concentrated in attribution rather than rubric coverage.
  • The authors disclose: “At CMU, Xinran Zhao is supported by the ONR Award N000142312840 and the Amazon AI PhD Fellowship.” Locator: Acknowledgments.
  • The authors state that training data and model checkpoints are open-sourced; the conclusion also states that code is open-sourced. Release contents are not verified in the supplied text.
  • The user-study ratings concern perceived usefulness and confidence; they do not establish objectively measured comprehension, citation accuracy or causal effects on business visibility. The reported subgroup means have no accompanying uncertainty estimates in this chunk.
  • Appendix A.3 says test-time intent awareness helps in most cases, not all; Table 7 includes declines, including llama-3.1-8B without training and its intent-explicit SFT variant.
  • Appendix A.8 reports a trade-off: flexible intent schemas can improve performance but produce inconsistent types across questions, whereas the fixed schema supports consistency and readability.
  • Footnote 1 narrows the intent scope: “We focus on the non-psycho-lingual functional intents and remove the emotion expressing mode.”
  • The source provides no direct measurement of brand inclusion, commercial discoverability or marketer credibility in AI answers.

Evidence behind this briefing

[c1] The experiments isolate writing rather than retrieval performance by fixing retrieved information per query. Evaluation covers SQA-CS-V2 (100 test and 100 validation samples), a modified DeepScholar task (63 papers, using title and ground-truth related-work subsection headers), and ResearchQA’s AI subset (50 test questions, without retrieval and with paragraph intents only). Official benchmark evaluations are used; SQA-CS-V2 uses LLM judges.

Section 4.1, Experimental Setting · Read the study

[c2] On SQA-CS-V2, intent prompting raises gemini-2.5-pro’s overall macro-average from 88.1 to 89.7 points, with unchanged rubric score and higher citation precision and recall. The reported paired-test p-values for overall improvement are 0.013 for Gemini and 0.072 for o3; the authors use alpha = 0.1, so o3 does not meet a 0.05 threshold.

Section 4.2, Table 2; columns Overall, Rubrics, Ans. P, Citation P, Citation R; significance discussion immediately below Table 3 · Read the study

[c3] For SQA-CS-V2, qwen3-8b intent-multiview SFT scores 88.6 overall versus 80.7 without training and 83.2 with baseline SFT. The intent-trained variants are also prompted to produce intents at inference. Training uses Gemini-generated reports for 1,000 OpenScholar queries; multiview expands each report into four instruction-report pairs, with training steps controlled for compute.

Section 4.2, Table 4; columns Overall, Rubrics, Answer P, Citation P, Citation R; training setup in Sections 3.4 and 4.1 · Read the study

[c4] Intent prompting is not uniformly beneficial: o3’s DeepScholar overall score falls from 46.8 to 43.2 with both intent types, while citation precision falls from 39.1 to 27.2 and claim coverage from 40.2 to 34.3. A qualitative inspection of only 20 o3 claims found added information from model memory in nearly 60% of cases. The reported paragraph-only recovery to 49.3 is attributed to an appendix not present here.

Section 4.2, o3 citation-behavior discussion; Table 3 · Read the study

[c5] A between-subject study of Gemini-generated reports finds higher self-reported navigational usefulness with intent annotations, not measured factual accuracy or objective comprehension. On 1–5 agreement scales, paragraph and citation ratings are 4.47±0.83 and 4.46±0.87 with intents versus 3.84±1.05 and 3.62±1.18 for baseline. The sample includes 20 participants, 71 reports, 349 unique paragraphs, and 416 unique citations; the chunk does not define the ± statistic or report significance.

Section 4.3, Case Study: Intents help navigate readers in model-generated long-form reports · Read the study

[c6] The authors explicitly limit demonstrated generalization to scientific writing: policy, law, and humanities may require intent categories absent from this schema. They also characterize the schema as purely synthetic rather than grounded in human writing-process observations.

Section 5, Generalization Across Domains; synthetic-schema qualification in The Complexity and Hierarchy of Intent · Read the study

[c7] Evaluation used official APIs for Gemini, Claude and GPT, locally served open models, and benchmark-specific LLM judges. Where applicable, generation used a 22,000-token maximum and temperature 1.0; otherwise original task settings were retained.

Appendix A.1 Implementation details · Read the study

[c8] The authors disclose LLM assistance for grammar correction and coding, but deny using LLMs to write the paper or construct the overall code-base logic.

Appendix A.2 The Use of Large Language Models (LLMs) · Read the study

[c9] The authors' no-risk assessment is scoped to long-form reports with human-verified queries; experiments were mainly in English, and datasets and experimental models were publicly available.

Appendix A.5 Ethical Statements · Read the study

[c10] On DeepScholar Bench, intent-aware inference degraded o3's citation quality, rather than universally improving it. The appendix therefore tested paragraph-intent-only inference and SFT variants under the main-paper settings without further training.

Appendix A.7 Extended DeepScholar Bench Results, discussion following Table 9 · Read the study

[c11] The user study measured self-reported reading guidance and confidence, not verified comprehension or answer accuracy. On its 1–5 agreement scale, the intent interface received higher paragraph/citation ratings than baseline in both recruitment groups: 4.26/4.19 versus 3.77/3.29 among 8 personally recruited participants, and 4.55/4.60 versus 3.94/4.06 among 12 Prolific participants.

Appendix A.9 User study details, User-centered task design; scale and question definitions in preceding study description · Read the study

[c12] The user-study recruitment targeted computer-science master's or PhD participants, used 20 recruits across personal advertising and Prolific, paid $30/hour, and specified exclusion rules for unusually scored or unusually fast annotations.

Appendix A.9 User study details, Participant pool · Read the study