The practical brief
The finding
Across 413 expert-authored U.S. legal research questions, the highest observed complete-answer pass rate was 42.9%, despite an 85.4% weighted partial-credit score. [c1] [c2] [c7] [c8]
Why it matters
A polished, citation-rich answer can still omit a required authority or get a consequential legal conclusion wrong.
| Measure | Observed result |
|---|---|
| Complete, source-verified answers | 42.9% of 413 question responses; ±2.4 percentage points standard error [c1] [c2] |
| Weighted partial credit | 85.4% mean share of available rubric weight across 413 question responses; ±1.1 percentage points standard error [c1] [c2] |
What to try
BLURSOR’s practical interpretation
Our interpretation: use legal research agents as inputs to licensed-attorney review, not as a substitute for it, and assess complete answers rather than presentation or search activity.
Study boundary
This fixed-snapshot benchmark does not certify deployment safety or measure business discoverability, brand inclusion, or marketing outcomes.
The decision: research assistance, not final legal authority
If your business uses AI to research legal obligations or prepare legally sensitive content, the risk is treating a plausible answer as a verified conclusion. This study supports keeping licensed-attorney review before reliance—not approving an answer because it reads confidently or includes links.
The researchers tested 13 frontier models on 413 open-ended U.S. legal research questions written and reviewed by attorneys. Agents could search the web and case law, fetch pages, and retrieve passages from collected documents. The evaluation therefore tested a tool-enabled research workflow, not merely a chatbot answering from memory.
Its primary measure was “all-pass”: every required grading criterion had to be satisfied, and cited authorities had to verify. An invalid cited URL made the whole response fail. Weighted pass, a separate partial-credit measure, tracked how much required content an answer covered.
High partial credit concealed incomplete answers
Claude Opus 4.8 had the highest observed all-pass rate: 42.9%, with a standard error of 2.4 percentage points. Its weighted pass score was 85.4%. These measures answer different questions: covering most required material is not the same as delivering a complete, grounded answer.
The shortfall was not simply cosmetic. Across evaluated responses, 29.5% missed at least one high-value criterion, such as the central legal conclusion; these misses accounted for roughly 41% of failures. Another 4.3% of runs met every content criterion but failed citation verification.
Tasks requiring reconciliation—combining multiple or conflicting legal authorities—were especially difficult. Pooled across models, their all-pass rate was 20.4%, versus 29.9% for questions without that attribute. This is an association, not proof that conflicting sources caused the failures.
A realistic interpretation: verify outcomes, not activity
More research activity did not reliably predict better answers across models. Higher turn counts, tool-call counts, and inference spending were not dependable signals of correctness. That does not mean reducing research effort improves accuracy.
Tool-enabled configurations outperformed no-tool, single-shot configurations, but tools did not make answers consistently reliable. Because that comparison changed both tool access and the interaction format, it cannot isolate the benefit of search alone.
For owners and marketers, the useful distinction is between an answer that looks supported and one whose required claims actually check out. The following are our workflow interpretations, not tactics tested by the study:
- Ask a reviewing attorney to check the conclusion, required authorities, and whether citations support the claims attributed to them.
- When comparing research systems, request complete-answer evaluations rather than relying only on partial-credit scores or search volume.
- Do not treat this benchmark as evidence that adding citations to your website will improve AI visibility or recommendations.
What the study cannot tell your business
The benchmark covers objectively answerable, English-language U.S. legal research using hypothetical facts. It excludes much legal strategy, counseling, negotiation, and drafting judgment. Its fixed model, tool, and authority snapshot cannot establish performance in your particular workflow.
Scoring depended on an AI judge calibrated against attorneys, rather than attorney review of every response. Each model ran once per question, so uncertainty estimates do not capture repeat-run variability. The highest observed score also does not establish a statistically significant lead over GPT-5.5.
The authors are affiliated with Vals AI, which maintains the associated live leaderboard. The supplied paper provides no separate funding or competing-interest declaration. It measures legal-answer reliability—not source-selection preferences, brand inclusion, or discoverability improvements.
The source and evidence
Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
Study limitations and disclosures
- Scope is limited to objectively answerable, English-language U.S. legal research at a fixed snapshot of models, tools, and authority. Results do not cover strategy, counseling, negotiation, or other tasks without a single correct answer, and do not measure marketing or discoverability outcomes.
- GPT-5.4 graded the benchmark. Calibration covered 341 rubric items from 45 responses to 15 questions, with three attorney raters; response-level all-pass agreement with the attorney majority was 86.7%. The authors describe the judge as mildly strict and scores as slightly conservative relative to human grading.
- Each model was evaluated once per question. Reported uncertainty reflects sampling across questions, not variation between repeated runs. The leaderboard supports performance tiers rather than a strict ordering; Opus 4.8's lead over GPT-5.5 was not statistically significant.
- Only five questions are publicly released; 200 validation questions and 208 test questions are withheld to limit contamination. Questions use original hypothetical facts rather than real client matters. Participating attorneys were compensated and consented to use of their contributions.
- All authors are affiliated with Vals AI in San Francisco. Vals AI maintains the associated live leaderboard, and the authors provide a public evaluation-harness repository. The supplied paper contains no separate funding or competing-interest declaration, so those disclosures are not available.
- Cross-model effort comparisons and reconciliation results are associations, not causal estimates. The no-tool comparison also switches to single-shot answering, preventing an isolated estimate of tool access alone.
- The authors explicitly state that this benchmark is not deployment certification or clearance for unsupervised legal use. Outputs require licensed-attorney review before reliance; fluent presentation can conceal an incorrect central conclusion.
Evidence behind this briefing
[c1] The benchmark measures complete, grounded correctness rather than plausibility or preference: every required rubric criterion must pass, and an invalid cited URL makes the entire response incorrect. Weighted pass is a separate partial-credit metric.
Section 3.3, Rubric and metrics · Read the study
[c2] Across all 413 questions, Claude Opus 4.8 had the highest observed all-pass rate, 42.9% ± 2.4 percentage points standard error, versus 85.4% ± 1.1 weighted pass. These results describe this tool-enabled test system, not deployment fitness; Appendix D does not establish a statistically significant lead over GPT-5.5.
Section 5.1, Table 3, Claude Opus 4.8 row; uncertainty and units defined in Table 3 caption; paired comparison in Appendix D, Table 7 · Read the study
[c3] Partial correctness can conceal substantive errors. Across the evaluated corpus and models, 29.5% of question responses missed at least one high-value criterion, accounting for roughly 41% of failures. Another 4.3% of runs passed every rubric criterion but failed citation verification.
Section 5.2, Partial correctness hides substantive omissions; Appendix E, Failure decomposition · Read the study
[c4] Pooled across the 13 models, reconciliation questions had lower all-pass accuracy: 20.4% (95% CI 16.1–24.9) versus 29.9% (26.4–33.4) without that attribute, a reported 9.4-percentage-point gap. There were 116 reconciliation questions among 413; labels overlap. Question-clustered intervals and adjusted regression support an association, not a causal effect of conflicting sources.
Section 5.3, Reliability varies systematically; denominator in Table 1; clustering and adjusted analysis in Appendix C · Read the study
[c5] Across models, more agent turns, tool calls, and inference spending did not reliably predict higher correctness. Opus averaged 11.7 turns with 42.9% all-pass, whereas Kimi averaged nearly 100 turns and 112 tool calls with 15.5%. This cross-model comparison does not establish that reducing research effort causes better answers.
Section 5.4, More turns, tool calls, and cost do not predict reliability · Read the study
[c6] In the full-corpus ablation, the tool-enabled configuration outperformed the no-tool, single-shot configuration for all 13 models. Mean all-pass fell by 22.5 percentage points and weighted pass by 25.6 points without tools. Because the comparison also changes the interaction setting to single-shot, it should not be treated as an isolated estimate of tool access alone.
Section 5.5, Table 4 caption · Read the study
[c7] Generalization is limited to U.S., English-language, objectively answerable legal research and a fixed model/tool/authority snapshot. Results depend on model-mediated grading, and one run per model per question means the reported uncertainty does not measure run-to-run variation.
Section 6, Limitations and future work · Read the study
[c8] The authors explicitly disclaim deployment certification and unsupervised legal use. Fluent, authoritative-looking output can still contain a wrong central conclusion; legal outputs require licensed-attorney review before reliance.
Ethics Statement, first paragraph · Read the study