The practical brief

The finding

Across four tasks and two generators, distributing 24 document slots across 12 responses increased portfolio recall by 14.4 percentage points versus one response containing 24 documents. [c1] [c7] [c12] [c18] [c20]

Why it matters

Businesses commissioning AI research tools face a coverage-versus-speed decision, not simply a choice about context size.

Coverage gains depend on evidence supply and the alternative
ComparisonReported result
12 × 2 documents versus 1 × 24 documents: primary pool+14.4 percentage points of portfolio recall across four tasks and two generators with a 400-document pool; 95% CI +11.9 to +17.0 points [c7] [c19]
Same allocation comparison: constrained-pool stress test+6.5 percentage points of portfolio recall on average across three decoding seeds with a 30-document pool [c20] [c21]
12-response portfolio versus output-token-matched structured answer+2.9 percentage points of portfolio recall at 24 documents under the integrate instruction; 95% CI +0.2 to +5.4 points; 480 paired queries [c18]
Sequential versus single-pass generation latency17.5 seconds versus 1.7 seconds in the reported deployment comparison; attribution probing adds nonzero overhead [c12]
Measured benchmark comparisons, not commercial AI-search results. The constrained-pool stress test is separate from the primary experiment; budget matching does not establish equal latency or total compute. Study source

What to try

BLURSOR’s practical interpretation

Our interpretation: compare repeated generation using fresh evidence against better-prompted, longer single answers, measuring coverage, accuracy, latency and total cost separately.

Study boundary

The outcome measures coverage across responses, not commercial AI visibility or single-answer accuracy. A separate limited-evidence stress test reduced the allocation advantage to 6.5 percentage points.

Decide whether broader coverage justifies a longer wait

If you are commissioning an AI research assistant, a bigger bundle of documents may not produce the most complete result. This study tests whether several responses using changing evidence cover more material than one response—and what that choice costs in time.

For marketers seeking inclusion in public AI answers, the implication is narrower: retrieved evidence and evidence actually used are different things. These experiments do not establish how to get a brand cited or recommended.

Evidence [c1] [c2] [c7] [c12]

The study measured coverage across multiple responses

The researchers studied retrieval-augmented generation, or RAG: supplying retrieved documents to a language model before it writes. Their main outcome was portfolio recall—the fraction of predefined correct answer units mentioned in at least one response. That is not the same as every response being accurate or useful.

The primary held-out evaluation used 100 queries per task across four tasks, three generators and two decoding seeds. Comparisons were paired and shared retrieval and generation settings. A separate allocation experiment crossed document counts and response counts across four tasks and two generators.

The authors also tested attribution: identifying which documents influenced generation. Their leave-one-out probe removes a document and checks the effect. In constructed pools with topical distractors, it identified reliance better than relevance scores.

Evidence [c1] [c2] [c7] [c14]

Fresh evidence and stronger single answers change the comparison

With 24 allocated document slots, two documents per response across 12 responses beat one response containing 24 documents by 14.4 percentage points of portfolio recall; the 95% confidence interval was 11.9–17.0 points. The primary factorial used a 400-document retrieval pool. In a separate three-seed stress test with only 30 documents, the same comparison averaged a 6.5-point advantage. This is a separate setting, not a replacement estimate or confidence interval: exhausting fresh evidence reduced the gain.

Rotating fresh evidence also beat repeatedly generating from the same context. Asking for all supported interpretations improved recall by 5.0–9.8 points; under that instruction, the 24-slot comparison still favored 12 responses by 11.1 points.

A structured single response received the same generated-output token allowance as a 12-response portfolio. It still trailed at every tested width and instruction, but the gap narrowed to 2.9 points with 24 documents and the broader instruction. Matching output tokens does not establish equal total compute.

The full proposed system combined feedback-driven evidence scheduling with steered decoding, which changes text generation. It beat the Carriage comparator by 3.3 points, but neither the full system nor scheduling alone significantly outperformed deep rotation. Wider contexts were not universally worse; rotation performed best in the recipe-adaptation task.

Evidence [c7] [c9] [c10] [c8] [c16] [c17] [c18] [c19] [c20] [c21]

Test simpler alternatives before custom orchestration

Our interpretation: iterative retrieval is a candidate for breadth-sensitive research, not a default for every customer interaction. Test better instructions and a longer structured answer before assuming custom scheduling is necessary.

The paper reports 17.5 seconds for sequential generation versus 1.7 seconds for the single-pass configuration, with additional attribution-probe overhead. Customer tolerance for that delay remains untested. Ask vendors or internal teams:

  • Is success measured across many outputs, or in the one answer a customer receives?
  • Is enough fresh, useful evidence available to sustain repeated retrieval?
  • Are latency and total costs measured with attribution probes included?
  • Is added coverage factually supported, rather than merely more extensive?

Evidence [c1] [c12] [c17] [c18] [c20]

The source and evidence

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

Peiyang Liu, Xi Wang, Di Liang and Wei Ye · 2026-08-24

Study limitations and disclosures

  • These controlled RAG experiments do not measure deployed commercial AI-search visibility, citations, recommendations or customer preference. Portfolio recall measures coverage across responses, not single-answer accuracy.
  • The allocation advantage depends on fresh evidence supply. The primary factorial used a 400-document pool; a separate three-seed stress test with 30 documents reduced the same allocation advantage to an average of 6.5 percentage points.
  • Equal document-slot or output-token budgets are not equal response-time or total-compute budgets. Attribution probes add overhead, and the hardware-specific cost frontier excludes probing.
  • The full proposed system did not significantly outperform deep rotation or its no-steering variant in the main held-out recall comparisons. Width effects varied, and the hierarchical analysis included only four task–generator conditions.
  • Attribution validation used specially selected passage pools, retaining approximately one-fifth of raw ASQA and one-third of QAMPARI. This limits its representativeness for unrestricted retrieval.

Evidence behind this briefing

[c1] The study evaluates portfolio recall: the fraction of gold answer units asserted in at least one generated response, rather than single-answer accuracy or user preference. Primary evaluation uses 100 queries per task across four tasks, three generators, and two decoding seeds, with paired inference and task-weighted bootstrap inference.

Sections 3.1 and 7.4, Evaluation Protocol and Zero-Leakage Rigor · Read the study

[c2] In constructed pools preserving two answer-bearing documents, topical relevance did not reliably identify evidence reliance when padding consisted of same-query passages entailing no gold answer aliases. At width 3, BM25 and query–document cosine had AUCs of 0.444 and 0.484, versus 0.876 for deletion leave-one-out; the latter achieved 0.824 at width 24. This is attribution discrimination, not answer accuracy.

Section 4.2, Deconstructing the Diagnostic Illusion via Controlled Pools · Read the study

[c3] The reported calibrated width elasticity is -0.68 (0.02), but it is not a universal coefficient: natural-pool redundancy strata yield slopes between -0.43 and -0.67. The result concerns thresholded attribution utilization under particular calibration protocols, not a direct measurement of all LLM attention or commercial search ranking.

Section 5.3, Empirical Law II; Figure 4 caption identifies the thresholded utilisation statistic · Read the study

[c4] The width–count grid compares ordinary generation using rotating rank windows against repeated sampling of a fixed top-k context, over four tasks and two generators. The reported equal-evidence-slot comparison—24 documents in one response versus two documents per response across 12 responses—improves portfolio recall by 0.144. Equal evidence slots do not establish equal latency or total compute; no uncertainty interval for this particular cross-configuration contrast appears in the supplied text.

Figure 1(a); Section 8.1 and Table 4 · Read the study

[c5] With the same ordinary sampler and the full custom decoder excluded, Ascp's evidence coverage rate is 0.626 versus 0.375 for deep rotation. However, deep rotation uses more documents and has slightly higher portfolio recall, 0.303 versus 0.300. This comparison supports extraction efficiency, not an unqualified claim of superior final recall.

Table 2, Policy dependence with one ordinary decoder; columns: policy, ECR, used, offered, PR@T · Read the study

[c6] Sequential generation carries higher autoregressive latency and probe overhead. The supplied introduction reports scaling verification up to 32B, but the corresponding later experimental results are not available in this chunk.

Section 1, Introduction · Read the study

[c7] In the reported factorial evaluation, allocating 24 document slots as 12 generations with two documents each rather than one generation with 24 documents increased portfolio recall by 0.144 absolute points (95% CI 0.119–0.170). This is a document-slot comparison, not an equal-latency comparison; the grid crossed four context widths and three generation counts over four tasks and two generators.

Section 8.1, budget-allocation comparison following Table 5 · Read the study

[c8] Width effects varied materially across task–generator conditions. For rotation at T=12, widening from k=2 to k=24 had a hierarchical grand mean of +0.025 portfolio-recall points and a prediction interval of −0.412 to +0.462. Only four conditions informed this analysis, limiting precision and contradicting a universal recommendation based on its mean alone.

Table 6, caption and T=12 column · Read the study

[c9] The matched fixed-context experiment separated repeated stochastic generation from exposure to fresh retrieved evidence. At 12 generations, rotation exceeded fixed-context resampling by reported absolute portfolio-recall margins of 0.087 at k=24 to 0.140 at k=2 (all p<0.001). These are coverage results, not measurements of answer accuracy or preference.

Section 8.2, matched fixed-context decomposition · Read the study

[c10] On held-out queries, full Ascp achieved equal-task portfolio recall of 0.309 versus Carriage's 0.276: +0.033 absolute points, 95% CI 0.022–0.044, BH q<0.001. Evaluation used 100 disjoint queries per task, four tasks, three generators and two decoding seeds, with paired comparisons and crossed bootstrap. Important null comparators remain: full Ascp versus deep rotation was +0.006 (q=0.343), and versus Ascp without steering was +0.009 (q=0.102); scheduler-only Ascp versus deep rotation was −0.003 (q=0.717).

Section 8.4; Tables 10 and 11 · Read the study

[c11] The separate five-seed decoder decomposition reported a full-decoder portfolio-recall gain of 0.0107 versus decoding off (seed-t4 CI 0.0072–0.0142, p=0.001), alongside +0.0555 ECR, +0.0275 lexical grounding and −1.0 words. Inference used equal task weights and a fixed three-model set. APC alone was null (−0.0009, p=0.775), and adding APC with contrast enabled reduced recall by 0.0080 (p=0.001); the supplied measurements do not establish a hallucination-prevention guarantee.

Table 12, caption and decoder-decomposition rows; Section 8.6 · Read the study

[c12] Document-slot matching did not eliminate deployment-cost differences. The text gives sequential-generation latency of 17.5 seconds versus 1.7 seconds for the single-pass (24,1) configuration and discloses additional nonzero teacher-forced LOO probing. Figure 6 measures Qwen/ASQA on an A100 with the probe removed; token cost favors narrow rotation while latency favors wide single-generation packaging, so its cost frontier is not a universal deployment prescription.

Section 8.7, deployment overhead; Figure 6 caption · Read the study

[c13] In 2,880 query-level observations across six context widths, evidence-utilization dilution persisted in low-redundancy strata. With pairwise entailment filtering, the utilized fraction fell from 0.575 at k=2 to 0.164 at k=24, with reported elasticity -0.528 (0.021). These are natural-pool, free-generation estimates, not the fixed-target protocol estimates from Section 5.3.

Appendix B.2, Multiplicity Control and Redundancy Stratification · Read the study

[c14] Benchmark comparisons shared the retriever, evidence pool, context size, portfolio size, and generation temperature. Architecture and facet hyperparameters were selected once on a disjoint 40-query ASQA development split and frozen before held-out algorithmic evaluations.

Appendix C, Reproducibility and Experimental Artifacts, opening protocol paragraph · Read the study

[c15] Attribution-probe validation used specially constructed ALCE passage pools rather than the full raw datasets. Eligibility required disjoint answer-bearing evidence and verified distractors, retaining approximately one-fifth of ASQA and one-third of QAMPARI; this qualifies the validation sample's representativeness.

Appendix C, controlled ground-truth validation-pool construction · Read the study

[c16] The recipe adaptation task is a material exception to a blanket preference for selective scheduling: rotation achieved peak portfolio extraction in this isolated, nearly flat-relevance setting. The task used 9,486 Spanish-origin recipes and rewarded valid Spanish ingredients attested in retrieved evidence but absent from the source recipe, not general answer accuracy.

Appendix D.1, Open-Domain Cross-Cultural Generalization · Read the study

[c17] An integrate instruction that requested all supported interpretations improved portfolio recall by +0.050 to +0.098 across the utility-side factorial, but did not eliminate the advantage of more generation rounds. The reported budget-matched comparison favored (k=2,T=12) over (k=24,T=1) by +0.111, with 95% CI [+0.087,+0.136]; Table 16 specifies 314–480 paired queries per cell.

Appendix D.2, Instructional Bounds and Structured Output Controls; Table 16 and following paragraph · Read the study

[c18] With the same decode-token budget as a T=12 portfolio, a structured single-response baseline had lower reported portfolio recall at every tested width and instruction, using 480 paired queries per row. Gaps narrowed with context width: under integrate at k=24, portfolio recall was 0.416 versus 0.388, a reported gap of +0.029 with 95% CI [+0.002,+0.054] and p<0.05.

Appendix D.2, Table 17 caption and integrate/k=24 row · Read the study

[c19] The primary width–count factorial used a 400-document retrieval pool, allowing up to 288 document slots without repeats. Its caption explicitly contrasts this with an earlier 30-document pool, where 12 rounds offered the same 24–30 documents.

Section 8.1, Table 4 caption · Read the study

[c20] The allocation advantage depends on the supply of fresh evidence. In a separate stress test using a constrained 30-document retrieval pool rather than the primary factorial's 400-document pool, the 2-documents × 12-responses versus 24-documents × 1-response comparison produced an average portfolio-recall advantage of 6.5 percentage points. This is a separate-setting result, not a replacement estimate or confidence interval for the primary 14.4-point gain.

Appendix B.1, discussion following Table 14 · Read the study

[c21] Across three decoding seeds in the separate 30-document-pool stress test, the budget-matched portfolio-recall advantages were 6.2, 6.1 and 7.2 percentage points, averaging 6.5 points with a cross-seed standard deviation of 0.58 percentage points. This standard deviation describes seed variation, not a confidence interval.

Appendix B.1, Table 14, budget (2,12) vs (24,1) row; columns seed 0, seed 1, seed 2, mean, s.d. · Read the study